I am trying to normalise between layers of my stacked LSTM network in PyTorch. The network looks something like this:
class LSTMClassifier(nn.Module):
def __init__(self, input_dim, hidden_dim, layer_dim, output_dim):
super().__init__()
self.hidden_dim = hidden_dim
self.layer_dim = layer_dim
self.lstm1 = nn.LSTM(input_dim, hidden_dim, layer_dim, batch_first=True)
self.lstm2 = nn.LSTM(input_dim, hidden_dim, hidden_dim, batch_first=True)
self.fc1 = nn.Linear(hidden_dim, 32)
self.fc2 = nn.Linear(32, 1)
self.dropout = nn.Dropout(p=0.2)
self.batch_normalisation1 = nn.BatchNorm1d(hidden_dim)
self.batch_normalisation2 = nn.BatchNorm1d(hidden_dim)
def forward(self, x):
h0, c0 = self.init_hidden(x)
out, (hn1, cn1) = self.lstm1(x, (h0, c0))
out = self.dropout(out) # error line
out = self.batch_normalisation1(out)
h1, c1 = self.init_hidden(out)
out, (hn2, cn2) = self.lstm2(out, (h1, c1))
out = self.dropout(out)
out = self.batch_normalisation1(out)
h2, c2 = self.init_hidden(out)
out, (hn3, cn3) = self.lstm2(out, (h2, c2))
out = self.dropout(out)
out = self.batch_normalisation1(out)
out = self.fc1(out[:, -1, :])
out = self.dropout(out)
out = self.fc2(out)
return out
def init_hidden(self, x):
h0 = torch.zeros(self.layer_dim, x.size(0), self.hidden_dim)
c0 = torch.zeros(self.layer_dim, x.size(0), self.hidden_dim)
return [t for t in (h0, c0)]
There is an error being raise where I have commented in the above, which is happening due to the BatchNorm1d expecting a 2 dimensional input.
In particular, I am initialising the model and passing batched data to the network as:
model = LSTMClassifier(5, 128, 3, 1)
model(X)
error:
RuntimeError: running_mean should contain 3 elements not 128
The X input tensor has shape torch.Size([10, 3, 5]), i.e. batch size is 10 and each input has dimensions 3 x 5 i.e. 5 features and 3 time steps.
The error is arising due to the BatchNorm1d trying to normalise across the wrong dimension - in the network the variable out has shape torch.Size([1, 3, 128]), i.e. the 5 input features are mapped to 128 hyper variables.
I could reshape the variable put inside the forward function, but this seems unnecessary. I have also tried using BatchNorm2d, but it requires a 4d Tensor, which my variable out is not. Is there some way to over come this?
Additionally, I am trying to add normalisation in my network to speed the training - I am not entirely sure how the PyTorch BatchNorm functions work, so would appreciate an explanation of this v. much. Specifically, why should we normalise across the time dimension and not the feature dimensions?