How can I implement multiple hidden layers in an RNN (PyTorch)?

Viewed 954

My Pytorch RNN for name classification does not allow me to choose multiple hidden layers. If I choose more than 1 layer I get the following error message:

Traceback (most recent call last): File "TRAIN2.py", line 63,

line 104, in training: y_hat = y_hat.view(batch_size, n_categories) RuntimeError: shape '[26718, 6]' is invalid for input of size 320616

I don't know exactly what the problem is but the input size is doubled when I choose 2 layers, tripled with 3 layers and so on.. My output size is 6 and when I choose 1 layer it works but with two layers there are twice as many inputs and I get the error and with 3 layers there are 3 times as many inputs. Shouldn't the input size stay the same regardless of the number of layers?

I am attaching the main file (where the error appears in the loss computation) and my network architecture.

for epoch in range(start_epoch, epochs):

   for batch_number, (x_batch, y_batch) in enumerate(batches):

      # Initialise the hidden layer
      hidden = model.initialise_hidden(batch_size=batch_size)

      if torch.cuda.is_available():
        x_batch = x_batch.to(device)
        y_batch = y_batch.to(device)

      # Model
      y_hat = rnn_gru(x_batch)

      # Loss
      y_hat = y_hat.view(batch_size, n_categories)
      loss = criterion(y_hat, y_batch)
      current_loss = current_loss + loss.item()

      # Restore Gradients
      optimizer.zero_grad()

      # Backward
      loss.backward()

      # Step
      optimizer.step()

And this is my model

class RNN_GRU(nn.Module):

  # input_size = Vocabulary size
  # hidden_size = Size of the hidden "units" in the GRU's
  # n_layers -> Connections between the GRU's
  # dropout -> fraction of nodes to be shut down

  def __init__(self, input_size, embedding_size, hidden_size, output_size, n_layers, dropout):
    super(RNN_GRU, self).__init__()

    # Members
    self.n_layers = n_layers
    self.hidden_size = hidden_size

    # 1) Embedding
    self.embedding = nn.Embedding(num_embeddings=input_size, embedding_dim=embedding_size)

    # 2) Packing
    #self.pack = nn.utils.rnn.pack_padded_sequence(batch_first=False)

    # 3) GRU
    self.gru = nn.GRU(input_size=embedding_size, hidden_size=hidden_size, num_layers=n_layers, bias=True, batch_first=False, dropout=dropout, bidirectional=False)

    # 4) Linear
    self.linear = nn.Linear(in_features=hidden_size, out_features=output_size, bias=True)

  def forward(self, x):

    # X -> [batch_size, n]
    #     [x11, x12, x13, ..., x1n1]
    # X:  [x21, x22, x23, ..., x2n2]
    #     [x31, x32, x33, ..., x3n3]
    # xYZ -> Name Y, character Z code
    # nY -> Length (num chars) of name Y. n1, n2, n3, ..., can be different (if no padding)

    batch_size = x.size(0)

    # 0.1) From the input name, get the embedding vector (for dimension requirements, it has to be transposed)
    input = self.embedding(x.t())

I am aware that using multiple layers does probably not increase performance but I want to use a non-zero dropout rate for which I'm being told that it is only active in all layers but the last and with only 1 layer therefore useless...

Thank you in advance for your help.

0 Answers
Related