Apologies for the poorly worded titled; I am new to NLP using Python and am having trouble. I have a dataframe of news articles that looks like this:
press =
| Article_Text | Date | Unigrams | Bigrams | Trigrams |
| -------------------- | ------ | ------------------------ | --------------------------------------- | ------------------------------------ |
| This is the Article | 5-5-21 | [this, is, the, article] | [(this, is), (is, the), (the, article)] | [(this, is, the),(is, the, article)] |
| A second Article | 9-2-20 | [a, second, article] | [(a,second), (second,article)] | [(a, second, article)] |
I want to take the top 30 most frequent uni-,bi-, and trigrams (30 total among all 3 categories, not 30 for each ngram) and then have those values become new columns in the dataframe. So a new column could be [this,is] if that was one of the most frequent occurring across all articles. The entries in each of these new fields would then be the tfidf score for that n-gram in that article (row). I'm able to get the current dataframe but cannot go further. My code:
def tokenize(article):
lemmatizer = nltk.stem.WordNetLemmatizer()
word_tokens = word_tokenize(article.lower())
clean_tokens = [lemmatizer.lemmatize(w) for w in word_tokens if not w in stop_words]
return clean_tokens
press['unigrams'] = press['Article_Text'].apply(tokenize)
press['bigrams'] = press['unigrams'].apply(lambda row: list(nltk.ngrams(row, 2)))
press['trigrams'] = press['unigrams'].apply(lambda row: list(nltk.ngrams(row, 3)))
I've tried using texthero tfidf and term frequency with no luck. Any point in the right direction would be appreciated.