Why are weight matrices shared between embedding layers in 'Attention is All You Need' paper?

Viewed 94

I am using the Transformer module in pytorch from the paper "Attention is All You Need". On page 5, the authors state that

In our model, we share the same weight matrix between the two embedding layers and the pre-softmax linear transformation, similar to [30]. (page 5)

The embedding layer, at least in pytorch, is a learnable tensor whose columns are the embedding vectors corresponding to each word. My confusion stems from the fact that in the paper, the Transformer learns a translation task between languages (i.e. English to German). Thus, how could the embedding weights be shared for the English and German embedding vectors?

In addition, how could the weights be shared between the output embedding (which goes from word index to embedding vector) and the linear layer (which goes from embedding vector to word probabilities)? As far as I can tell there is no constraint requiring the embedding tensor must be orthogonal (so that its inverse is its transpose).

1 Answers

Encoder and Decoder have different tokenizers and token embeddings, one for the source language, one of the target language. The shared weights belong to embedding layer of the decoder (the target language) and softmax layer of the decoder (again, the target language), hence it is the same language.

Assume that vocabulary size V = 32_000 and embedding size E = 768. Then the weights of the embedding layer is of shape V x E. Consequently, the last layer of the decoder will have a weight matrix of shape H x V, where H is hidden dimension for that layer. If you set H equal to E, so that E = V, then you can transpose the embedding weight matrix V x E into E x V, which allows you to re-use it before Softmax activation. This is how they can be shared.

Related