These two attentions are used in seq2seq modules. The two different attentions are introduced as multiplicative and additive attentions in this TensorFlow documentation. What is the difference?
These two attentions are used in seq2seq modules. The two different attentions are introduced as multiplicative and additive attentions in this TensorFlow documentation. What is the difference?
I just wanted to add a picture for a better understanding to the @shamane-siriwardhana
the main difference is in the output of the decoder network
There are actually many differences besides the scoring and the local/global attention. A brief summary of the differences:
The good news is that most are superficial changes. Attention as a concept is so powerful that any basic implementation suffices. There are 2 things that seem to matter though - the passing of attentional vectors to the next time step and the concept of local attention(esp if resources are constrained). The rest dont influence the output in a big way.
For more specific details, please refer https://towardsdatascience.com/create-your-own-custom-attention-layer-understand-all-flavours-2201b5e8be9e
Luong-style attention: scores = tf.matmul(query, key, transpose_b=True)
Bahdanau-style attention: scores = tf.reduce_sum(tf.tanh(query + value), axis=-1)