LSTM based Text Generation with Tensorflow : Problem With Generating Output Text

Viewed 81

I'm trying to make an LSTM based text generator with Tensorflow. But I've some confusion regarding the sensible output generation. I'm trying to describe the whole situation concisely.

My Training Dataset: Bangla Song Lyrics (4000+ lyrics)

By taking two subsequent lines I've prepared a dataset with 30,000 samples.

  • Minimum Sample Length: 25
  • Maximum Sample Length: 120

I've tokenized each sample into characters and prepared a vocab dictionary that contains 181 unique vocabs.

I've prepared a char_id dictionary to encode each sample which contains ID 1 to 181. Each sample is encoded with these ID and 0 is used for padding.

Here is my model definition:

model = Sequential([
    Embedding(vocab_size+1,64),
    LSTM(512,return_sequences=True),
    LSTM(512,return_sequences=True),
    Dense(vocab_size, activation='softmax')

model.compile(optimizer=tf.keras.optimizers.Adam(),
              loss = [tf.keras.losses.SparseCategoricalCrossentropy(from_logits=True)],
              metrics=['accuracy']
             )
])

Now here my embedding input dimension is vocab_size+1 which is according to the documentation. And the last layer dimension is vocab_size since I want the model to predict a character with ID between 1 to 181.

Here is my train and target sample:

train = [ 1  2  3  4  5  6  7  8  3  4  5  9 10  3 11 12 13  7  4 14  9  3 15 16
  5 17  8 18  5 19 20  4  3 21 18  6 14  0  0  0  0  0  0  0  0  0  0  0
  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0
  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0
  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0]

target = [ 2  3  4  5  6  7  8  3  4  5  9 10  3 11 12 13  7  4 14  9  3 15 16  5
 17  8 18  5 19 20  4  3 21 18  6 14  0  0  0  0  0  0  0  0  0  0  0  0
  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0
  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0
  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0  0]

You can see that I have shifted the target one index right to let the model predict the next character.

Now for generating the predicted output I use np.argmax() by passing the returned array from model.prediction(tensor). The np.argmax() returns the position with the highest value and I'm taking it as an ID and converting it to the character. Am I doing it correctly?

Do I have to try np.argmax(pred) + 1? Since I have vocab set with ID 1 to 181. But predicted array has indices from 0 to 180.

Another problem is the validation loss becomes nan after the few steps in the first epoch and it remains nan. Setting clip value didn't solve the problem. I've seen that if the last layer dimension is set to vocab_size+1 it fixes the problem. But this is also confusing.

Till now, I couldn't generate any text that makes sense. I'm trying hard to figure out what I am missing. I've described my current understanding too, so that, anyone can help me to understand my misunderstanding.

0 Answers
Related