I am following a tutorial on how to build an ngram model (https://www.kaggle.com/alvations/n-gram-language-model-with-nltk) and I would like to adjust the code so that I am able to remove stop words and punctuation after tokenising the text. However, when I try, no stopwords are removed (tokenized_text and clean_text contain the same items, nothing is removed). Can someone help me figure out what I did wrong/provide alternative solutions? Thanks a lot!
text = gutenberg.raw('carroll-alice.txt')[1:600]
# tokenise and turn everything into lowcase
tokenized_text = [list(map(str.lower, word_tokenize(sent)))
for sent in sent_tokenize(text)]
print(tokenized_text)
clean_text = [word for word in tokenized_text if word not in stopwords.words('english')]
print(clean_text)