I'm actually trying to make a rudimentary classifier so I'd be OK with an NLTK solution but my first couple of attempts at this have been with Pandas.
I have a couple of lists that I want to check the text for and get word counts on, and then return an ordered
import pandas as pd
import re
fruit_sentences = ["Monday: Yellow makes me happy. So I eat a long, sweet fruit with a peel.",
"Tuesday: A fruit round red fruit with a green leaf a day keeps the doctor away.",
"Wednesday: The stout, sweet green fruit keeps me on my toes!",
"Thursday: Another day with the red round fruit. I like to keep the green leaf.",
"Friday: Long yellow fruit day, peel it and it's ready to go."]
df = pd.DataFrame(fruit_sentences, columns = ['text'])
banana_words = ['yellow', 'long', 'peel']
apple_words = ['round', 'red', 'green leaf']
pear_words = ['stout', 'sweet', 'green']
print(df['text'].str.count(r'[XYZ_word in word list]'))
Here is where the code blows up because the str.count() does not accept a list.
The end goal is to have a returned list of tuples like this:
fruits = [('banana', 5), ('pear', 6), ('apple', 6)]
Yes, I could iterate over all the lists to do this but it seems like I just don't know enough Python rather than Python doesn't know how to handle this elegantly.
I found this question but it looks like everyone answered it incorrectly or with a different solution than what was actually asked for, it's here.
Thank you for helping this newbie figure it out!