Pytorch,How to feed output of CNN into input of RNN?

Viewed 3133

I am new to CNN, RNN and deep learning. I am trying to make architecture that will combine CNN and RNN. input image size = [20,3,48,48] a CNN output size = [20,64,48,48] and now i want cnn ouput to be RNN input but as I know the input of RNN must be 3-dimension only which is [seq_len, batch, input_size] How can I make 4-dimensional [20,64,48,48] tensor into 3-dimensional for RNN input?

and another question how do I initiate the first hidden state with

torch.zeros()

I don't know what exact information I should pass in this function. the only thing that I know it is

[layer_dim, batch, hidden_dim]

Thank you.

3 Answers

I assume that 20 here is size of a batch. In that case, set batch = 20.

seq_len is the number of time steps in each stream. Since one image is input at one time step, seq_len = 1.

Now, 20 images of size (64, 48, 48) has to be converted for the format

Since the size of input is (64, 48, 48), input_size = 64 * 48 * 48

model = nn.LSTM(input_size=64*48*48, hidden_size=1).to(device)

#Generating input - 20 images of size (60, 48, 48)
cnn_out = torch.randn((20, 64, 48, 48)).requires_grad_(True).to(device)

#To pass it to LSTM, input must be of the from (seq_len, batch, input_size)
cnn_out = cnn_out.view(1, 20, 64*48*48)

model(cnn_out)

This will give you the result.

By following @Arun soulution. Finally I can pass image tensor though RNN Layer But the problem after is some how pytorch need the first hidden state as [1, 1, 1] only. I don't know why. And now my output of RNN is [1, 20, 1] . I thought my output will be [1, 20, 147456] . So I can reshape my output shape back to Image input shape [20, 64, 48, 48]

class Rnn(nn.Module):
    def __init__(self):
        super(Rnn, self).__init__()
        self.rnn = nn.RNN(64*48*48, 1, 1, batch_first=True, nonlinearity='relu')

    def forward(self, x):
        batch_size = x.size(0)
        hidden = self.init_hidden(batch_size)
        images = x.view(1, 20, 64*48*48)
        out, hidden = self.rnn(images, hidden)
        out = torch.reshape(out, (20,64,96,96))
        return out, hidden

    def init_hidden(self, batch_size):
        hidden = torch.zeros(1, 1, 1).to(device)
        return hidden

Your question is very interesting.The output of CNN is 4 dimension,but the input of RNN require 3 dimension.

Obviourly, you know the meaning of dimension. The problem is samply be shape operation.

Related