I am trying to replicate my code from Keras into PyTorch to compare the performance of multi-layer bidirectional LSTM/GRU models on CPUs and GPUs. I would like to look into different merge modes such as 'concat' (which is the default mode in PyTorch), sum, mul, average. Merge mode defines how the output from the forward and backward direction will be passed on to the next layer.
In Keras, it's just an argument change for the merge mode for a multi-layer bidirectional LSTM/GRU models, does something similar exist in PyTorch as well? One option is to do the merge mode operation manually after every layer and pass to next layer, but I want to study the performance, so I want to know if there is any other efficient way.
Thanks,