I am using IndicBART for sequence-to-sequence prediction on tamil sentences.
I have trained the model on 100,000 samples of tamil data for 30 epochs, then tested predictions on a few sentences. These are the predictions given by the model- I have presented both the token IDs output by the model and the tokens decoded from the token IDs:
source: 'bullet es'
predicted: ''
predicted token IDs: [64000, 64000, 64000, 64000, 64000, 64000, 64000, 64000, 64000, 64000, 64000, 64000, 64000, 64000, 64000, 64000, 64000, 64000, 64000, 2]
gold: 'bullet e s'
source: 'இங்க பழச த்துலேந்து பேசுறங்க'
predicted: ''
predicted token IDs: [64000, 64000, 64000, 64000, 64000, 64000, 64000, 64000, 64000, 64000, 64000, 64000, 64000, 64000, 64000, 64000, 64000, 64000, 64000, 2]
gold: 'இங்க palladam துல இருந்து பேசறங்க'
And so on. All the predictions are blank lines.
Untrained IndicBART is not producing blank predictions, only the trained checkpoint.
Is this a feature of IndicBART, or am I doing something wrong?