I am trying to count the number of keywords from a pandas DataFrame as such:
df = pd.read_csv('amazon_baby.csv')
selected_words = ['awesome', 'great', 'fantastic', 'amazing', 'love', 'horrible', 'bad', 'terrible', 'awful', 'wow', 'hate']
The selected_words have to be counted from the Series: df['review']
i have tried
def word_counter(sent):
a={}
for word in selected_words:
a[word] = sent.count(word)
return a
and then
df['totalwords'] = df.review.str.split()
df['word_count'] = df.totalwords.apply(word_counter)
----------------------------------------------------------------------------
----> 1 df['word_count'] = df.totalwords.apply(word_counter)
c:\users\admin\appdata\local\programs\python\python36\lib\site-packages\pandas\core\series.py in apply(self, func, convert_dtype, args, **kwds)
3192 else:
3193 values = self.astype(object).values
-> 3194 mapped = lib.map_infer(values, f, convert=convert_dtype)
3195
3196 if len(mapped) and isinstance(mapped[0], Series):
pandas/_libs/src\inference.pyx in pandas._libs.lib.map_infer()
<ipython-input-51-cd11c5eb1f40> in word_counter(sent)
2 a={}
3 for word in selected_words:
----> 4 a[word] = sent.count(word)
5 return a
AttributeError: 'float' object has no attribute 'count'
can someone help..? i am guessing it is because of some fault value in the series that is not a string. . .
some people have tried helping but the issue is that the individual cells in the DataFrame have sentences in them.
I need to extract a count of selected words, preferably in dictionary form and store them in a new column in the same dataFrame with the corresponding rows.
