My Pytorch RNN for name classification does not allow me to choose multiple hidden layers. If I choose more than 1 layer I get the following error message:
Traceback (most recent call last): File "TRAIN2.py", line 63,
line 104, in training: y_hat = y_hat.view(batch_size, n_categories) RuntimeError: shape '[26718, 6]' is invalid for input of size 320616
I don't know exactly what the problem is but the input size is doubled when I choose 2 layers, tripled with 3 layers and so on.. My output size is 6 and when I choose 1 layer it works but with two layers there are twice as many inputs and I get the error and with 3 layers there are 3 times as many inputs. Shouldn't the input size stay the same regardless of the number of layers?
I am attaching the main file (where the error appears in the loss computation) and my network architecture.
for epoch in range(start_epoch, epochs):
for batch_number, (x_batch, y_batch) in enumerate(batches):
# Initialise the hidden layer
hidden = model.initialise_hidden(batch_size=batch_size)
if torch.cuda.is_available():
x_batch = x_batch.to(device)
y_batch = y_batch.to(device)
# Model
y_hat = rnn_gru(x_batch)
# Loss
y_hat = y_hat.view(batch_size, n_categories)
loss = criterion(y_hat, y_batch)
current_loss = current_loss + loss.item()
# Restore Gradients
optimizer.zero_grad()
# Backward
loss.backward()
# Step
optimizer.step()
And this is my model
class RNN_GRU(nn.Module):
# input_size = Vocabulary size
# hidden_size = Size of the hidden "units" in the GRU's
# n_layers -> Connections between the GRU's
# dropout -> fraction of nodes to be shut down
def __init__(self, input_size, embedding_size, hidden_size, output_size, n_layers, dropout):
super(RNN_GRU, self).__init__()
# Members
self.n_layers = n_layers
self.hidden_size = hidden_size
# 1) Embedding
self.embedding = nn.Embedding(num_embeddings=input_size, embedding_dim=embedding_size)
# 2) Packing
#self.pack = nn.utils.rnn.pack_padded_sequence(batch_first=False)
# 3) GRU
self.gru = nn.GRU(input_size=embedding_size, hidden_size=hidden_size, num_layers=n_layers, bias=True, batch_first=False, dropout=dropout, bidirectional=False)
# 4) Linear
self.linear = nn.Linear(in_features=hidden_size, out_features=output_size, bias=True)
def forward(self, x):
# X -> [batch_size, n]
# [x11, x12, x13, ..., x1n1]
# X: [x21, x22, x23, ..., x2n2]
# [x31, x32, x33, ..., x3n3]
# xYZ -> Name Y, character Z code
# nY -> Length (num chars) of name Y. n1, n2, n3, ..., can be different (if no padding)
batch_size = x.size(0)
# 0.1) From the input name, get the embedding vector (for dimension requirements, it has to be transposed)
input = self.embedding(x.t())
I am aware that using multiple layers does probably not increase performance but I want to use a non-zero dropout rate for which I'm being told that it is only active in all layers but the last and with only 1 layer therefore useless...
Thank you in advance for your help.