Trying to recreate BiLSTM model from Adhikari et al. 2019 (LSTM_reg) in Tensorflow

Viewed 40

Im trying to recreate this model LSTM_reg from this paper in TensorFlow to use in my problem. I've come up with the following code:

def get_model(lr=0.001):
    model = tf.keras.models.Sequential()
    model.add(tf.keras.layers.Embedding(nb_words, output_dim=embed_size, weights=[embedding_matrix], input_length = maxlen, trainable=False))
    model.add(tf.keras.layers.Dropout(0.2)) # embedding dropouts
    model.add(tf.keras.layers.Bidirectional(tf.keras.layers.LSTM(256, return_sequences=True, recurrent_dropout=0.2, activation = 'tanh'))) # weight drop on recurrent layers using recurrent_dropout
    model.add(tf.keras.layers.MaxPooling1D(pool_size=2, padding = 'valid'))
    model.add(tf.keras.layers.Flatten())
    model.add(tf.keras.layers.Dense(512, activation='relu'))
    model.add(tf.keras.layers.Dropout(0.5))
    model.add(tf.keras.layers.Dense(20))
    model.add(tf.keras.layers.Activation('sigmoid'))

    model.compile(loss = 'categorical_crossentropy' , optimizer = 'adam', metrics = ['accuracy', tfa.metrics.F1Score(num_classes = 20)])
    return model

Have I gone about this the right way? Got some pretty weird values while training my dataset, hence was wondering about my implementation.. There is a pytorch implementation for this model here. But I'm not sure if I have reproduced this correctly.

1 Answers

One major difference I can see is that the paper uses a global max pooling, whereas you've only used max pooling with a kernel size of 2:

def get_model(lr=0.001):
    model = tf.keras.models.Sequential()
    model.add(tf.keras.layers.Embedding(nb_words, output_dim=embed_size, weights=[embedding_matrix], input_length = maxlen, trainable=False))
    model.add(tf.keras.layers.Dropout(0.2)) # embedding dropouts
    model.add(tf.keras.layers.Bidirectional(tf.keras.layers.LSTM(256, return_sequences=True, recurrent_dropout=0.2, activation = 'tanh'))) # weight drop on recurrent layers using recurrent_dropout
    model.add(tf.keras.layers.GlobalMaxPooling1D())
    model.add(tf.keras.layers.Dense(512, activation='relu'))
    model.add(tf.keras.layers.Dropout(0.5))
    model.add(tf.keras.layers.Dense(20))
    model.add(tf.keras.layers.Activation('sigmoid'))

    model.compile(loss = 'categorical_crossentropy' , optimizer = 'adam', metrics = ['accuracy', tfa.metrics.F1Score(num_classes = 20)])
    return model

I obviously don't have your data so make sure that the arguments are set correctly: https://keras.io/api/layers/pooling_layers/global_max_pooling1d/.

Another change is that the pytorch repo you shared has a ReLU after the LSTM (I don't know why). You could try adding that in and see if it helps.

Related