PyTorch Adding src_key_padding_mask in TransformerEncoder leads to inf loss

Viewed 1024

I'm using this code as base to build an own transformer model, using different inputs as presented there. The part in class TransformerModel(nn.Module): or in the following my own implementation shows some problems:

def make_len_mask(self, inp):
    return (inp == 0).transpose(0, 1)


class TransformerModel(nn.Module):
    def __init__(self):
        encoder_layer = TransformerEncoderLayer(ninp, nhead, nhid, dropout)
        encoder_norm = LayerNorm(ninp)
        self.encoder = TransformerEncoder(encoder_layer, nlayers, encoder_norm)

    def forward(self, src, trg):
        src.shape # (x, y, z)
        trg.shape # (x, y)
        # eliminate last dimension of source tensor, which is (batch_size, samples, features) to compute mask
        # resulting in a [true, false]-Vector indicating which elements are padding elements
        padding_tensor = src.mean(2) # padding_tensor.shape: (x,y)
        src_pad_mask = self.make_len_mask(padding_tensor)
        # self.src.mask = None
        output = self.encoder(src, mask=self.src_mask, src_key_padding_mask=src_pad_mask)

Using that src_pad_mask results in a ValueError: The loss returned in training_step is nan or inf.. If not using that mask in EncoderLayer there are results.

My inputs are

model(source, target)

where source are continuous and target are words structured as follows:

target = [1] + [4,2,3,8] + [99]  # 0 and 99 are start and end of sentence tokens

I tried to remove them with target_tensor[target_tensor == 1] = 0 or target_tensor[target_tensor == 99] = 0 this unfortunately lead to a RuntimeError: one of the variables needed for gradient computation has been modified by an inplace operation: ... so I'm providing the sequence with sos and eos tokens, which doesnt feel right, maybe the problem arises there? But the sequence is already padded there so its not possible to remove the last or first index as suggested in the example.

If using nn.Transformer() instead of single EncoderLayer () the results are strongly overfitting without a mask or the same error arises using said mask. Using just a target_mask doesnt produce correct inputs.

Is there any possibility to find out where this error arises or if my mask is calculated wrongly? More discussion on github. Is it necessary to provide masks during training or inference? If so Im not doing it, maybe someone could help or point to a source if its true?

# Values != 0 => False
# Values == 0 => True
src_pad_mask:
tensor([[False],
        [False],
        ...
        [ True],
        [ True]], device='cuda:0')
0 Answers
Related