I tried to see if the sentence contains any words in the keywords (fuzzy match). It works well in the single line but when I run through my dataframe, it returns []
def word2vec(word):
from collections import Counter
from math import sqrt
# count the characters in word
cw = Counter(word)
# precomputes a set of the different characters
sw = set(cw)
# precomputes the "length" of the word vector
lw = sqrt(sum(c*c for c in cw.values()))
# return a tuple
return cw, sw, lw
def cosdis(v1, v2):
# which characters are common to the two words?
common = v1[1].intersection(v2[1])
# by definition of cosine distance we have
return sum(v1[0][ch]*v2[0][ch] for ch in common)/v1[2]/v2[2]
list_of_keywords = ['Clearing', 'grubbing']
Sentence = 'Clear and grub Low density Light vegetation'
def list1(Sentence, list_of_keywords):
a = [x for x in Sentence.str.split() for y in list_of_keywords if cosdis(word2vec(x), word2vec(y)) > 0.7]
return a
This is a code I used for function:
df_fill = df_fill.astype('str')
df_fill['lol']=df_fill.apply(lambda x: list1(x['Line Item Description'],x['Keyword']), axis=1)
Is there any faster way to do it as my dataframe has 10000 rows? For this input:
list_of_keywords = ['Clearing', 'grubbing']
Sentence = 'Clear and grub Low density Light vegetation'
The expected output is:
['Clear','grub']
Added in a new column.
Example dataframe:
Sentence | keyword | Result
Clear and grub Low density Light vegetation | ['Clearing', 'grubbing'] | ['Clear','grub']