I want to use an implementation of an attention mechanism by Yang et al.. I found a working implementation of a custom layer that uses this attention machanism here. Instead of using the output values of my LSTM:
my_lstm = LSTM(128, input_shape=(a, b), return_sequences=True)
my_lstm = AttentionWithContext()(my_lstm)
out = Dense(2, activation='softmax')(my_lstm)
I would like to use the hidden states of the LSTM:
my_lstm = LSTM(128, input_shape=(a, b), return_state=True)
my_lstm = AttentionWithContext()(my_lstm)
out = Dense(2, activation='softmax')(my_lstm)
But I get the error:
TypeError: can only concatenate tuple (not "int") to tuple
I tried it in combination with return_sequences but everything I've tried failed so far. How can I modify the returning tensors in order to use it like the returned output sequences?
Thanks!