I am quite confused on how I can build and use an N-gram model using NLTK in Python. I was going through the documentation and wanted to create a trigram model based on a simple corpus below.
from nltk import ngrams
from nltk.lm.preprocessing import pad_both_ends
from nltk.lm.preprocessing import flatten
from nltk.lm import MLE
n=3
corpus = [
'natural language processing is a subfield of linguistics computer science and artificial intelligence concerned with the interactions between computers and human language in particular how to program computers to process and analyze large amounts of natural language data',
'speech recognition is an interdisciplinary subfield of computer science and computational linguistics that develops methodologies and technologies that enable the recognition and translation of spoken language into text by computers with the main benefit of searchability',
'natural language generation is a software process that produces natural language output while it is widely agreed that the output of any nlg process is text there is some disagreement on whether the inputs of an nlg system need to be nonlinguistic'
]
padded_corpus = [list(pad_both_ends(s.split(), n=n)) for s in corpus]
ngram_list = [list(ngrams(text, n=n)) for text in padded_corpus]
vocab = list(flatten(pad_both_ends(s.split(), n=n) for s in corpus))
lm = MLE(n)
lm.fit(ngram_list, vocab)
print("Vocab length:", len(lm.vocab))
Until here, everything runs fine. However, when I would like to experiment with the trained model, such as returning the probability scores of bigrams:
lm.score("natural", ["language"])
lm.score("software", ["product"])
lm.score("is", ["a"])
they all return a probability of 0. And when attempting to generate text:
lm.generate(3, random_seed=5)
it returns the following error message:
ValueError: Can't choose from empty population
which I don't understand, since lm.vocab is clearly non-empty, the model seems to have learned the vocabulary.