I have a dataframe df that has a column containing text df['text'] (articles from a newspaper, in this case). How can I get a count of the rows in df['text'] that have a word count above some threshold of n words?
An example of df is shown below. Each article can contain an arbitrary number of words.
print(df['text'].head(10))
0 Emerging evidence that Mexico economy was back...
1 Chrysler Corp Tuesday announced million in new...
2 CompuServe Corp Tuesday reported surprisingly ...
3 CompuServe Corp Tuesday reported surprisingly ...
4 If dining at Planet Hollywood made you feel li...
5 Hog prices fell Tuesday after government slaug...
6 Blue chip stocks rallied Tuesday after the Fed...
7 Sprint Corp Tuesday announced plans to offer I...
8 Shoppers are loading up this year on perennial...
9 Kansas and Arizona filed lawsuits against some...
Name: text, dtype: object
My goal with this data is to find a count of the articles that contain greater than n words. See the psuedocode below for an example.
n = 250 # number of words cutoff for counting
counter = 0
for row in df['text']:
if df['text'].wordcount >= n: # wordcount is some function on a df that counts the words in a string for one row
counter += 1
print(counter)
The desired output is number of articles containing more than n words (in this case, n is arbitrarily set to 250). So, in the psuedocode above, wordcount is some function that counts the words in one row (or, in this case, a single article). Thus for row x if N (number of words in the article) is 340, it would be greater than n, which is set at a threshold of 250. Therefore the if statement would be triggered and counter would increase by one.
Ideally, I would like to do this in a vectorized way, as the dataframe is large. If not, apply works just fine.