I have a column in pandas dataframe with millions of rows. Many words are non-English (e.g. words from other languages or that do not mean anything, like "**5hjh"). I thought of using Wordnet as a comprehensive English dictionary to help me clean up this column, which comprises lists. Ideally, the output should be a new column with English words only.
I have tried the following code, which I got from Stackoverflow, but it does not seem to be working as it returns an empty column with no words whatsoever:
from nltk.corpus import wordnet
def check_for_word(s):
return ' '.join(w for w in str(s).split(',') if len(wordnet.synsets(w)) > 0)
df["new_column"] = df["original_column"].apply(check_for_word)