# Error using tensorMmul on permuted tensor

**URL:** <https://community.konduit.ai/t/error-using-tensormmul-on-permuted-tensor/675>\
**Category:** SameDiff\
**Created:** [July 1, 2020, 4:18pm UTC](https://community.konduit.ai/t/error-using-tensormmul-on-permuted-tensor/675 "2020-07-01T16:18:49Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![Dris101](https://avatars.discourse-cdn.com/v4/letter/d/94ad74/32.png) [@Dris101](https://community.konduit.ai/u/Dris101)\
**Post date:** [July 1, 2020, 4:18pm UTC](https://community.konduit.ai/t/error-using-tensormmul-on-permuted-tensor/675/1 "2020-07-01T16:18:49Z")

</div>

Hi there. I’m building a SameDiffVertex and have run into an issue where calling tensorMmul on a tensor which has previously been permuted causes a NPE in reduce.TensorMul.doDiff when calculating gradients

java.lang.NullPointerException: org.nd4j.linalg.api.ops.impl.reduce.TensorMmul.doDiff(TensorMmul.java:146)

which seems to be:-

int aAxes = range(0, larg().getShape().length);

I have tried a test case where I just use an sd.constant as the first arg to tensorMmul and try it with and without permuting it first, and the error goes away if I don’t do the permute. Unfortunately this is not an option in my real world case. Is it possible these two ops don’t play well together or am I doing something dumb somewhere?

---

<div class="post-metadata">

**Author:** ![treo](https://yyz1.discourse-cdn.com/flex035/user_avatar/community.konduit.ai/treo/32/47_2.png) [@treo](https://community.konduit.ai/u/treo)\
**Post date:** [July 1, 2020, 6:04pm UTC](https://community.konduit.ai/t/error-using-tensormmul-on-permuted-tensor/675/2 "2020-07-01T18:04:31Z")

</div>

It looks like that operation is still using an old java based backprop implementation, and it may be that there is a problem.

Without the code that you’ve tried it is hard to tell what exactly is going on though.

---

<div class="post-metadata">

**Author:** ![Dris101](https://avatars.discourse-cdn.com/v4/letter/d/94ad74/32.png) [@Dris101](https://community.konduit.ai/u/Dris101)\
**Post date:** [July 1, 2020, 6:51pm UTC](https://community.konduit.ai/t/error-using-tensormmul-on-permuted-tensor/675/3 "2020-07-01T18:51:23Z")

</div>

Thanks for quick response. Understood. It’s probably a showstopper for me unless there is a workaround. Let me try to pull a minimal gist together (will be in scala if that’s OK) since it looks like it might be a bug, but in brief if I have (with weight an appropriately shaped matrix):-

val testInputNotPermuted = sd.constant(“test\_input\_not”, Nd4j.ones(5, 3, 4, 7))  
sd.tensorMmul(“output”, testInputNotPermuted, weight, Array(3), Array(0), false, false, false)

the grads will calc using sd.calculateGradients. But if I do:-

val testInputPermuted = sd.constant(“test\_input”, Nd4j.ones(5, 7, 4, 3)).permute(0, 3, 2, 1)  
sd.tensorMmul(“output”, testInputPermuted, weight, Array(3), Array(0), false, false, false)

I will get the NPE when I call calculateGradients. The forward pass seems correct in both cases however.

---

<div class="post-metadata">

**Author:** ![treo](https://yyz1.discourse-cdn.com/flex035/user_avatar/community.konduit.ai/treo/32/47_2.png) [@treo](https://community.konduit.ai/u/treo)\
**Post date:** [July 1, 2020, 7:34pm UTC](https://community.konduit.ai/t/error-using-tensormmul-on-permuted-tensor/675/4 "2020-07-01T19:34:26Z")

</div>

It is possible to work around it in a similar way to how the workaround in this case worked:

> [@How to use sd.nn.batchNorm(…) in Deeplearning4j?](https://community.konduit.ai/t/how-to-use-sd-nn-batchnorm-in-deeplearning4j/85/6):
>
> Here’s one possible workaround: class BatchNormFixed(sameDiff: SameDiff?, input: SDVariable?, mean: SDVariable?, variance: SDVariable?, gamma: SDVariable?, beta: SDVariable?, epsilon: Double, vararg axis: Int) : org.nd4j.linalg.api.ops.impl.layers.convolution.BatchNorm(sameDiff, input, mean, variance, gamma, beta, epsilon, axis) { override fun doDiff(f1: MutableList\<SDVariable\>?): MutableList\<SDVariable\> { var list = args().toMutableList() list.add(f1!!.get(0)) retur…

However, you’ve got to understand how the native op works. If I get the opportunity, I’ll write up a workaround for this case too.

---

<div class="post-metadata">

**Author:** ![Dris101](https://avatars.discourse-cdn.com/v4/letter/d/94ad74/32.png) [@Dris101](https://community.konduit.ai/u/Dris101)\
**Post date:** [July 1, 2020, 10:05pm UTC](https://community.konduit.ai/t/error-using-tensormmul-on-permuted-tensor/675/5 "2020-07-01T22:05:44Z")

</div>

That would be great! Is the issue likely to be with sd.permute or sd.tensorMmul? I am guessing it’s tensorMmul as I’m pretty sure I’ve have tried mmul (rather than tensorMmul) on permuted tensors and it has been fine. Somewhat beyond my current understanding of the codebase unfortunately.

Would you like me to create an issue for this or wait until you have had a further look?

---

<div class="post-metadata">

**Author:** ![Dris101](https://avatars.discourse-cdn.com/v4/letter/d/94ad74/32.png) [@Dris101](https://community.konduit.ai/u/Dris101)\
**Post date:** [July 2, 2020, 2:25pm UTC](https://community.konduit.ai/t/error-using-tensormmul-on-permuted-tensor/675/6 "2020-07-02T14:25:42Z")

</div>

> <https://gist.github.com/Dris101/c2663600f3d35733cef2b9331d424ec7>

---

<div class="post-metadata">

**Author:** ![Dris101](https://avatars.discourse-cdn.com/v4/letter/d/94ad74/32.png) [@Dris101](https://community.konduit.ai/u/Dris101)\
**Post date:** [July 20, 2020, 11:17am UTC](https://community.konduit.ai/t/error-using-tensormmul-on-permuted-tensor/675/7 "2020-07-20T11:17:01Z")

</div>

I’ve been taking a look at this and as you suggested have overridden the _doDiff_ in _TensorMmul_ to use the  
existing C++ implementation in _libnd4j/include/ops/declarable/generic/blas/tensormmul.cpp_.  
This has been a partial success. The issue I have now is that, from what I can see, the implementation of _CUSTOM\_OP\_IMPL(tensormmul\_bp)_ is not working for all cases. Specifically it will generally error out if  
the ranks of the the two input tensors differ. The forward pass is fine however in those cases. From  
adding some extra _nd4j\_verbose_ calls in my own fork of the libnd4j code it appears that the exception is being raised by shapeUtils when called from the section

_// calculate dLdA_

_MmulHelper::tensorDot(dLdC, B, dLdA, axesBdLdC, axesB, permutAt);_

A _“ShapeUtils::evalShapeForTensorDot method: the numbers of a axes and b axes to make dot product along must have identical values !”_ runtime error will be thrown at this point.

To test further I used TensorFlow from python to generate a series of random test cases for input tensors A and B of different ranks, making sure that they were sized so that I could contract over at least one index, and dumping the resulting tensordot output and the A,B grads to .npy files and loading them back into ND4J and comparing the results with what I was getting from my overridden _tensormMul_ op which uses the C++ implementation for both the forward and backward passes. What I have so far, taking all combinations of tensor ranks for A and B from 2 to 6 is:-

Output from ScalaTest. The input shapes are in the square brackets and I have contracted over the second last index in all cases (though that is easy to change or randomise)

**Pass**  
[info] a=[2,1], b=[2,1] Calc pass Grad pass  
[info] a=[2,1], b=[3,2,1] Calc pass Grad pass  
[info] a=[1,4,4], b=[2,4,4] Calc pass Grad pass  
[info] a=[1,4,4], b=[1,4,4,4] Calc pass Grad pass  
[info] a=[2,3,2,2], b=[1,3,2,2] Calc pass Grad pass  
[info] a=[2,3,2,2], b=[1,1,4,2,2] Calc pass Grad pass  
[info] a=[4,4,4,4,2], b=[4,1,3,4,2] Calc pass Grad pass  
[info] a=[4,4,4,4,2], b=[2,2,2,1,4,2] Calc pass Grad pass  
[info] a=[3,3,2,1,4,3], b=[4,2,4,2,4,3] Calc pass Grad pass

**Fail**  
[info] a=[2,1], b=[2,3,2,1] Calc pass Grad crash  
[info] a=[2,1], b=[1,2,1,2,1] Calc pass Grad crash  
[info] a=[2,1], b=[3,3,1,1,2,1] Calc pass Grad crash  
[info] a=[1,4,4], b=[4,4] Calc pass Grad crash  
[info] a=[1,4,4], b=[1,1,1,4,4] Calc pass Grad crash  
[info] a=[1,4,4], b=[3,3,4,4,4,4] Calc pass Grad crash  
[info] a=[2,3,2,2], b=[2,2] Calc pass Grad crash  
[info] a=[2,3,2,2], b=[2,2,2] Calc pass Grad crash  
[info] a=[2,3,2,2], b=[2,2,4,3,2,2] Calc pass Grad crash  
[info] a=[4,4,4,4,2], b=[4,2] Calc pass Grad crash  
[info] a=[4,4,4,4,2], b=[2,4,2] Calc pass Grad crash  
[info] a=[4,4,4,4,2], b=[3,2,4,2] Calc pass Grad crash  
[info] a=[3,3,2,1,4,3], b=[4,3] Calc pass Grad crash  
[info] a=[3,3,2,1,4,3], b=[2,4,3] Calc pass Grad crash  
[info] a=[3,3,2,1,4,3], b=[2,4,4,3] Calc pass Grad crash  
[info] a=[3,3,2,1,4,3], b=[3,3,3,4,3] Calc pass Grad crash

The good news is that the forward pass (“Calc”) passes for all cases. The grads agree when they don’t error out as described above, but for most cases the grads will not calculate. All the cases where A and B are of the same rank do work, but this should not be a requirement in the general case. I note that all the tests of _tensormul\_bp_ in  
_libnd4j/tests\_cpu/layers\_tests/DeclarableOpsTests15.cpp_  
only consider cases where the rank of A and B are the same, so this would not have been picked up by those tests.

If this is genuinely an issue and I haven’t got the wrong idea, let me know what would be useful. I can raise an issue, put some code up in a gist etc. Given that TensorFlow implements a dense layer in terms of tensordot (I believe), not being able to backprop it would seem to be potentially problematic, not just for my somewhat esoteric use case.

---

<div class="post-metadata">

**Author:** ![agibsonccc](https://yyz1.discourse-cdn.com/flex035/user_avatar/community.konduit.ai/agibsonccc/32/697_2.png) [@agibsonccc](https://community.konduit.ai/u/agibsonccc)\
**Post date:** [July 21, 2020, 12:28am UTC](https://community.konduit.ai/t/error-using-tensormmul-on-permuted-tensor/675/8 "2020-07-21T00:28:12Z")

</div>

> [@Dris101](#):
>
> s. The issue I have now is that, from what I can see, the implementation of _CUSTOM\_OP\_IMPL(tensormmul\_bp)_ is not working for all cases. Specifically it will generally error out if  
> the ranks of the the two input tensors differ. The forward pass is fine however in those cases. From  
> adding some extra _nd4j\_verbose_ calls in my own fork of the libnd4j code it appears that the exception is being raised by shapeUtils when called from the section
> 
> _// calculate dLdA_
> 
> _MmulHelper::tensorDot(dLdC,_

Hi, yes please file an issue. Thanks!

---

<div class="post-metadata">

**Author:** ![Dris101](https://avatars.discourse-cdn.com/v4/letter/d/94ad74/32.png) [@Dris101](https://community.konduit.ai/u/Dris101)\
**Post date:** [July 21, 2020, 9:28am UTC](https://community.konduit.ai/t/error-using-tensormmul-on-permuted-tensor/675/9 "2020-07-21T09:28:03Z")

</div>

That’s done. Gist at:-

> <https://gist.github.com/Dris101/9983f0aedaa78d9477b4d3a1c8770bd0>
