# Tuning a simple 2 layer LSTM

**URL:** <https://community.konduit.ai/t/tuning-a-simple-2-layer-lstm/160>\
**Category:** Tuning Help\
**Created:** [February 22, 2020, 6:32pm UTC](https://community.konduit.ai/t/tuning-a-simple-2-layer-lstm/160 "2020-02-22T18:32:19Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![nimishatandon](https://yyz1.discourse-cdn.com/flex035/user_avatar/community.konduit.ai/nimishatandon/32/74_2.png) [@nimishatandon](https://community.konduit.ai/u/nimishatandon)\
**Post date:** [February 22, 2020, 6:32pm UTC](https://community.konduit.ai/t/tuning-a-simple-2-layer-lstm/160/1 "2020-02-22T18:32:19Z")

</div>

Hi am training an LSTM model with 2 layers and a batch size of 32 on a data set of 15000 chat utterances however it’s taking almost 4 hours to train. The stranger thing is that it takes the same amount of time on my laptop which is an i7 4 core and 8 logical processor vs that on a Linux bare metal server which has 256 GB ram and 16 cores. Any insights would be helpful here.  
I did go through ur link on performance tuning and tried some suggestions there but nothing seems to be working.

---

<div class="post-metadata">

**Author:** ![treo](https://yyz1.discourse-cdn.com/flex035/user_avatar/community.konduit.ai/treo/32/47_2.png) [@treo](https://community.konduit.ai/u/treo)\
**Post date:** [February 23, 2020, 7:43pm UTC](https://community.konduit.ai/t/tuning-a-simple-2-layer-lstm/160/2 "2020-02-23T19:43:56Z")

</div>

I’ve split this from the introductory post. In a separate topic it is a whole lot easier to talk about your problem.

> [@nimishatandon](#):
>
> The stranger thing is that it takes the same amount of time on my laptop which is an i7 4 core and 8 logical processor vs that on a Linux bare metal server which has 256 GB ram and 16 cores. Any insights would be helpful here.

That isn’t strange at all. Your description sounds like you have a rather small network and you are using a recurrent network. This means that most of the computations have to be done in sequence. More parallel resources don’t help in that case.

However, in order for us to help you here, we will need to know a bit more about the model you are training. Can you share your training code here?

---

<div class="post-metadata">

**Author:** ![nimishatandon](https://yyz1.discourse-cdn.com/flex035/user_avatar/community.konduit.ai/nimishatandon/32/74_2.png) [@nimishatandon](https://community.konduit.ai/u/nimishatandon)\
**Post date:** [February 23, 2020, 9:33pm UTC](https://community.konduit.ai/t/tuning-a-simple-2-layer-lstm/160/3 "2020-02-23T21:33:17Z")

</div>

Hi Treo, unfortunately I cannot paste the code exactly here. But I’ll try and give details as mush as I can here.

Building a ComputaionalGraph with a configuration using :  
SGD as the optimization algorithm  
Adam as the updated using l2 regularization and  
Xavier as the weight initializer.  
L1 with an input feature vector for each token in an utterance with 100 hidden layers followed by an Rnnoutput layer  
Activation sigmoid and loss as binary crossentropy.

The model is trained with a batch size of 32 over a number of epochs until the stopping criteria is met.

I have 2 CPUs and 32 cores wondering if there is anything at all that I could potentially parallelize to train my model faster.

---

<div class="post-metadata">

**Author:** ![treo](https://yyz1.discourse-cdn.com/flex035/user_avatar/community.konduit.ai/treo/32/47_2.png) [@treo](https://community.konduit.ai/u/treo)\
**Post date:** [February 24, 2020, 6:53am UTC](https://community.konduit.ai/t/tuning-a-simple-2-layer-lstm/160/4 "2020-02-24T06:53:35Z")

</div>

Not having the code makes things a bit harder to quantify.

But from what you’ve given us so far, it looks like you have a very small network and your batch size is also quite small. So there just isn’t that much computation to to parallelize during a single batch.

What you can do however, is try to use `ParallelWrapper` to train on multiple batches at once. It basically uses your computer the same way that a spark cluster would be used.

In our examples it is always used in conjunction with a multi-gpu setup, but it doesn’t require it. So if you have multiple cpu’s or if your network is too small to effectively use the single cpu you have, then you _may_ get a faster training with ParallelWrapper.

See [ParallelWrapper](https://javadoc.io/static/org.deeplearning4j/deeplearning4j-parallel-wrapper/1.0.0-beta6/index.html?org/deeplearning4j/parallelism/ParallelWrapper.html) for the JavaDoc (beta6). And you will have to add another dependency to your project:

```xml
 <dependency>
    <groupId>org.deeplearning4j</groupId>
    <artifactId>deeplearning4j-parallel-wrapper</artifactId>
    <version>1.0.0-beta6</version>
</dependency>

```

**But,** before you go down this route, you should probably also add a [PerformanceListener](https://javadoc.io/doc/org.deeplearning4j/deeplearning4j-nn/latest/org/deeplearning4j/optimize/listeners/PerformanceListener.html) to your model, to see where most of the time is spent. It could very well be that most of the time is spent in ETL (i.e. loading your data) instead of training the model.

---

<div class="post-metadata">

**Author:** ![nimishatandon](https://yyz1.discourse-cdn.com/flex035/user_avatar/community.konduit.ai/nimishatandon/32/74_2.png) [@nimishatandon](https://community.konduit.ai/u/nimishatandon)\
**Post date:** [February 24, 2020, 8:19am UTC](https://community.konduit.ai/t/tuning-a-simple-2-layer-lstm/160/5 "2020-02-24T08:19:41Z")

</div>

Thanks Treo will try that. I have already checked the ETL using the Performance listener and it consistently gives 0 except for the first batch . So don’t think ETL is the bottle neck.

Another thing, using AVX2 with beta6 did seem to decrease the time by 10% (just looking at time for few epochs) still need to try that on a bigger set for the full train.

Another question I had is to do with OMP\_NUM\_THREADS parameter , is that something I should be explicitly setting ?

Bw I m exposing the model building as a web service which may trigger separate threads for separate sets of data to use for model training .

---

<div class="post-metadata">

**Author:** ![treo](https://yyz1.discourse-cdn.com/flex035/user_avatar/community.konduit.ai/treo/32/47_2.png) [@treo](https://community.konduit.ai/u/treo)\
**Post date:** [February 24, 2020, 9:09am UTC](https://community.konduit.ai/t/tuning-a-simple-2-layer-lstm/160/6 "2020-02-24T09:09:37Z")

</div>

If you have an AVX2 or AVX512 capable processor, using the appropriate packages should help speed up the computation too. But, they too will only get you that far.

> [@nimishatandon](#):
>
> OMP\_NUM\_THREADS

We typically use half of the reported threads of the system, because using hyperthreading usually leads to reduced performance on numerical code. If you are training multiple models at once, reducing that number might be beneficial though. What exactly works on your system is something only you can figure out by trying different values.

---

<div class="post-metadata">

**Author:** ![nimishatandon](https://yyz1.discourse-cdn.com/flex035/user_avatar/community.konduit.ai/nimishatandon/32/74_2.png) [@nimishatandon](https://community.konduit.ai/u/nimishatandon)\
**Post date:** [March 2, 2020, 7:02pm UTC](https://community.konduit.ai/t/tuning-a-simple-2-layer-lstm/160/7 "2020-03-02T19:02:58Z")

</div>

Thanks @treo ! Over the last week I have been experimenting with a lot of combinations , training 1 model at a time using AVX2 with and without mil. Also using parallelwrapper .  
I am seeing some positive results. However when I use parallel wrapper I don’t see consistent results. And am guessing the reason for that could be averaging and the way it selecting the batch of data. Could you please clarify if there is a way to get reproducible results when the data set does not change ?

---

<div class="post-metadata">

**Author:** ![treo](https://yyz1.discourse-cdn.com/flex035/user_avatar/community.konduit.ai/treo/32/47_2.png) [@treo](https://community.konduit.ai/u/treo)\
**Post date:** [March 2, 2020, 7:37pm UTC](https://community.konduit.ai/t/tuning-a-simple-2-layer-lstm/160/8 "2020-03-02T19:37:35Z")

</div>

If I remember correctly, when training with parallel wrapper, you have something like a cluster running locally. The calculations aren’t run in lock-step. Which means that on each re-run, each thread can be done sooner or later, and thereby get updates at different points in time.

@raver119 can it be configured to behave deterministically?
