Pytorch Softmax giving nans and negative values as output

Viewed 3946

I am using softmax at the end of my model.

However after some training softmax is giving negative probability.In some situations I have encountered nans as probability as well.

one solution i found on searching is to use normalized softmax…however I can not find any pytorch imlpementaion for this.

Can someone please help to let know if there is a normalized softmax available or how to achieve this so that forward and backward propagations are smooth.

Please note that I am already using torch.nn.utils.clip_grad_norm_(model.parameters(), 40) to avoid exploding gradients

I am using pytorch 1.6.0

1 Answers

Softmax will always return positive results, but it will keep track of other results:

m = nn.Softmax(dim=1)
input = torch.randn(2, 3)
print(input)
output = m(input)
output

Out:

tensor([[ 0.0983,  0.4150, -1.1342],
        [ 0.3411,  0.5553,  0.0182]])

tensor([[0.3754, 0.5152, 0.1094],
        [0.3375, 0.4181, 0.2444]])

You are tracking the rows. Note how for

0.0983, 0.4150, -1.1342 You will get 0.3411, 0.5553, 0.0182

Saying that 0.4150 is the biggest value.

The hard max (as we know this is max()) will just return the maximum value.

So, if you have negative results for softmax this is not possible, you may have hit some implementation failure.

Related