Pytorch LSTM derivatives of output w.r.t. inputs are wrong

Viewed 91

I'm using LSTMs (Long Short Term Memory Neural Network) for learning certain time series, but the application is not really important here.

When computing the automatic derivatives of a Pytorch LSTM outputs w.r.t. to its inputs, I obtain wrong results for part of the inputs range (compared to numerical derivatives).

Moreover, when passing a subset of the original inputs, the outputs don't agree with the original ones. Since I'm new to recurrent neural networks, I suppose I'm lacking some understanding there and would be grateful, if someone could explain this unexpected behaviour.

The behaviour can be reproduced with this minimal script using a simple LSTM with one input and one output:

import torch
from torch import nn
from torch import autograd
import matplotlib.pyplot as plt

### USER INPUT
input_dim = 1
hidden_dim = 20
num_layers = 1
output_dim = 1

# PyTorch random number generator
torch.manual_seed(1234)


### DEFINITIONS
class Lstm(nn.Module):
    """Simple Long Short Term Memory Neural Network."""
    def __init__(self, input_dim: int, hidden_dim: int, num_layers: int, output_dim: int):
        """Initialise simple LSTM.

        Parameters
        ----------
        input_dim : int
            Input dimesion of NN
        hidden_dim : int
            Dimension of hidden layers of NN.
        num_layers : int
            Number of hidded layers in NN.
        output_dim : int
            Output dimesion of NN
        """
        # call __init__ from parent class
        super().__init__()

        # store inputs
        self.input_dim = input_dim
        self.hidden_dim = hidden_dim
        self.num_layers = num_layers

        # recurrent neural net
        self.lstm = nn.LSTM(input_dim, hidden_dim, num_layers, batch_first=True)

        # Fully connected layer
        self.fc = nn.Linear(hidden_dim, output_dim)

    def forward(self, input: torch.Tensor) -> torch.Tensor:
        """Forward pass through LSTM NN.

        Parameters
        ----------
        input : torch.Tensor
            Inputs to NN.

        Returns
        -------
        torch.Tensor
            Outputs of NN.
        """

        t = input.float().unsqueeze(0)

        # initialise hidden layer and cell state
        batch_size = 1  # use only a single batch for simplicity
        hidden = torch.zeros(self.num_layers, batch_size, self.hidden_dim)
        cell = torch.zeros(self.num_layers, batch_size, self.hidden_dim)

        # Passing in the input and hidden state into the model and obtaining outputs
        out, hidden = self.lstm(t, (hidden, cell))

        # Reshaping the outputs such that it can be fit into the fully connected layer
        out = out.squeeze()
        out = self.fc(out)

        return out


### COMPUTATION
# create LSTM with one input and one output dimension
my_lstm = Lstm(input_dim, hidden_dim, num_layers, output_dim)

# prepare two overlapping time series
times1 = torch.linspace(0.0, 1.0, 100).unsqueeze(1)
times2 = torch.linspace(0.2, 1.0, 80).unsqueeze(1)
times1.requires_grad = True
times2.requires_grad = True

# evaluate neural net with times
outputs1 = my_lstm.forward(times1)
outputs2 = my_lstm.forward(times2)

# compute automatic derivatives of first time series
auto_deriv1 = autograd.grad(outputs1,
                            times1,
                            torch.ones(times1.shape),
                            retain_graph=True,
                            create_graph=True)[0]

# compute numerical derivatives of first time series
numeric_deriv1 = (outputs1[1:-1] - outputs1[0:-2]) / (times1[1:-1] - times1[0:-2])

# plot outputs
plt.figure()
plt.plot(times1.detach().numpy(), outputs1.detach().numpy(), 'b')
plt.plot(times2.detach().numpy(), outputs2.detach().numpy(), 'r')
plt.grid("on")
plt.xlabel("LSTM Input")
plt.ylabel("LSTM Output")
plt.legend(["Time Series", "Time Series Late Start"])
plt.savefig("lstm_output.png")

# plot derivatives of first time series
plt.figure()
plt.plot(times1.detach().numpy(), auto_deriv1.detach().numpy(), 'b')
plt.plot(times1.detach().numpy()[0:-2], numeric_deriv1.detach().numpy(), 'b--')
plt.grid("on")
plt.xlabel("LSTM Input")
plt.ylabel("LSTM Output Derivative")
plt.legend(["Automatic", "Numerical"])
plt.savefig("lstm_derivatives.png")

Here are the plots I obtain.

  1. LSTM outputs of both time series. I would expect no deviation between both curves. LSTM outputs of both time series.

  2. Derivatives (automatic vs. numerical) of LSTM output w.r.t. inputs. I would expect both curves to be identical. Derivatives (automatic vs. numerical) of LSTM output w.r.t. inputs.

Since the deviations are in particular at the start of the time series, my suspicion would be that it has to do with the recurrent nature of the neural net used here. However, I don't really understand why identical input values should not provide identical outputs and in particular why the automatic derivatives are wrong.

0 Answers
Related