How to take most frequent n-grams across all dataframe rows and transpose as columns, with new entries being the tf-idf score?

Viewed 99

Apologies for the poorly worded titled; I am new to NLP using Python and am having trouble. I have a dataframe of news articles that looks like this:

press = 

    | Article_Text         | Date   | Unigrams                 | Bigrams                                 | Trigrams                             |
    | -------------------- | ------ | ------------------------ | --------------------------------------- | ------------------------------------ |
    | This is the Article  | 5-5-21 | [this, is, the, article] | [(this, is), (is, the), (the, article)] | [(this, is, the),(is, the, article)] |   
    | A second Article     | 9-2-20 | [a, second, article]     | [(a,second), (second,article)]          | [(a, second, article)]               |

I want to take the top 30 most frequent uni-,bi-, and trigrams (30 total among all 3 categories, not 30 for each ngram) and then have those values become new columns in the dataframe. So a new column could be [this,is] if that was one of the most frequent occurring across all articles. The entries in each of these new fields would then be the tfidf score for that n-gram in that article (row). I'm able to get the current dataframe but cannot go further. My code:

def tokenize(article):
    lemmatizer = nltk.stem.WordNetLemmatizer()
    word_tokens = word_tokenize(article.lower())  
    clean_tokens = [lemmatizer.lemmatize(w) for w in word_tokens if not w in stop_words]
    return clean_tokens

press['unigrams'] = press['Article_Text'].apply(tokenize)
press['bigrams'] = press['unigrams'].apply(lambda row: list(nltk.ngrams(row, 2)))
press['trigrams'] = press['unigrams'].apply(lambda row: list(nltk.ngrams(row, 3)))

I've tried using texthero tfidf and term frequency with no luck. Any point in the right direction would be appreciated.

0 Answers
Related