The accepted answer of this question How to add attention layer to a Bi-LSTM gives two options as -
- Attention with return_sequence=True
- Attention with return_sequence=False
May please someone let me know what may be the difference In the interpretation of the output of both the cases .. What is output represents in case 1 or in case 2
I am trying to implement a Custom Attention layer that can give a weighted sum of context vector at each time step for example if Dim(h)=256 and T=1,2,...128 I am trying to implement an Attention layer such that for each time step it will give a weighted sum that is a output of Dim(128,256).
Any help is highly appreciated Thanks in Advance.