# GPU Performance is worse than CPU

**URL:** <https://community.konduit.ai/t/gpu-performance-is-worse-than-cpu/3070>\
**Category:** Tuning Help\
**Created:** [February 9, 2024, 12:57pm UTC](https://community.konduit.ai/t/gpu-performance-is-worse-than-cpu/3070 "2024-02-09T12:57:19Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![baedaron](https://yyz1.discourse-cdn.com/flex035/user_avatar/community.konduit.ai/baedaron/32/1132_2.png) [@baedaron](https://community.konduit.ai/u/baedaron)\
**Post date:** [February 9, 2024, 12:57pm UTC](https://community.konduit.ai/t/gpu-performance-is-worse-than-cpu/3070/1 "2024-02-09T12:57:19Z")

</div>

I am training 2 Auto-Encoders.  
One on CPU and the other on GPU.  
With same training data types.

See below 2 logs(iteration time intervals).

OS: Windows 11  
DL4J Version: M2.1

Please help!

CPU LOG:

 ![CPU](https://canada1.discourse-cdn.com/flex035/uploads/konduit/original/2X/8/8a97da6518230ad456a961de43f29bee66eec0d6.png)

GPU LOG:

 ![GPU](https://canada1.discourse-cdn.com/flex035/uploads/konduit/original/2X/4/44063c39a7f4a1914eb13a1b29915f949f5769d2.png)

---

<div class="post-metadata">

**Author:** ![treo](https://yyz1.discourse-cdn.com/flex035/user_avatar/community.konduit.ai/treo/32/47_2.png) [@treo](https://community.konduit.ai/u/treo)\
**Post date:** [February 10, 2024, 2:05pm UTC](https://community.konduit.ai/t/gpu-performance-is-worse-than-cpu/3070/2 "2024-02-10T14:05:43Z")

</div>

This usually happens when you the GPU is used very inefficiently. For example, if you have a very small batch size or if you have a very small model. In those cases the transfer latency between the host system and the GPU dominates the overall time taken.

Can you share some more information about how big your model is?

---

<div class="post-metadata">

**Author:** ![baedaron](https://yyz1.discourse-cdn.com/flex035/user_avatar/community.konduit.ai/baedaron/32/1132_2.png) [@baedaron](https://community.konduit.ai/u/baedaron)\
**Post date:** [February 10, 2024, 5:13pm UTC](https://community.konduit.ai/t/gpu-performance-is-worse-than-cpu/3070/3 "2024-02-10T17:13:08Z")

</div>

thanks for your response.

below is more information.  
do you need more?

* * *

```
DataSetIterator cachedDsIter = new ExistingMiniBatchDataSetIterator(cacheDir, "dataset-%d.bin");
    // dataset-%d.bin: 168KB
// totalMiniBatchCount: 18074
// miniBatchSize: 64
   AsyncDataSetIterator asyncDsIter = new AsyncDataSetIterator(cachedDsIter, 64);

```

* * *

## . . .

```
double learningRatePositive = 0.001D;
int featureCount = 64;

    MultiLayerConfiguration config = new NeuralNetConfiguration.Builder()
		.trainingWorkspaceMode(WorkspaceMode.ENABLED)
		.inferenceWorkspaceMode(WorkspaceMode.ENABLED)
		.cacheMode(CacheMode.DEVICE)

        .seed(seed)
        .optimizationAlgo(OptimizationAlgorithm.STOCHASTIC_GRADIENT_DESCENT)
        .weightInit(WeightInit.XAVIER)
        .activation(Activation.TANH)
        .updater(Nadam.builder().learningRate(learningRate).build())
        .l2(0.0005)
        .list()
        .layer(0, new LSTM.Builder().activation(Activation.TANH).nIn(featureCount).nOut(featureCount / 4).build())
        .layer(1, new LSTM.Builder().activation(Activation.TANH).nIn(featureCount / 4).nOut(featureCount / 4 / 4).build())
        .layer(2, new LSTM.Builder().activation(Activation.TANH).nIn(featureCount / 4 / 4).nOut(featureCount / 4 / 4 / 4).build())
        .layer(3, new LSTM.Builder().activation(Activation.TANH).nIn(featureCount / 4 / 4 / 4).nOut(featureCount / 4 / 4).build())
        .layer(4, new LSTM.Builder().activation(Activation.TANH).nIn(featureCount / 4 / 4).nOut(featureCount / 4).build())
        .layer(5, new RnnOutputLayer.Builder().activation(Activation.IDENTITY).nIn(featureCount / 4).nOut(featureCount).lossFunction(LossFunction.MSE).build())
        .build();

```

* * *

 ![TaskManager](https://canada1.discourse-cdn.com/flex035/uploads/konduit/original/2X/1/1ff4024cb6b6da69e4149986634a1c9a344824c5.png)

---

<div class="post-metadata">

**Author:** ![treo](https://yyz1.discourse-cdn.com/flex035/user_avatar/community.konduit.ai/treo/32/47_2.png) [@treo](https://community.konduit.ai/u/treo)\
**Post date:** [February 15, 2024, 4:31am UTC](https://community.konduit.ai/t/gpu-performance-is-worse-than-cpu/3070/4 "2024-02-15T04:31:51Z")

</div>

LSTMs tend to amplify the problem I talked about in my last post.

You’ve got a mini batch size of 64 entries and 64 features with just a few layers, that have to work mostly sequentially.

That means your gpu is doing essentially no work between waiting for data to move.

You could try a larger batch size, given the comments in the code you’ve posted here, you can probably fit your entire dataset into gpu memory at once.

---

<div class="post-metadata">

**Author:** ![baedaron](https://yyz1.discourse-cdn.com/flex035/user_avatar/community.konduit.ai/baedaron/32/1132_2.png) [@baedaron](https://community.konduit.ai/u/baedaron)\
**Post date:** [February 15, 2024, 10:45pm UTC](https://community.konduit.ai/t/gpu-performance-is-worse-than-cpu/3070/5 "2024-02-15T22:45:28Z")

</div>

ok, i will try.  
thanks.
