I have a dataframe that looks like but it is around 1000 rows:
ID Messages
1 I eat apples
2 I eat oranges
3 I like to ski
4 I do not like vegetables.
5 The sky is blue.
I want to calculate the word mover's distance between all messages. My code currently looks like:
import gensim.downloader as api
#importing word vectors
wiki_vectors = api.load('glove-wiki-gigaword-50')
#initializing word2vec model
base_model = Word2Vec(vector_size = 50, min_count= 5)
base_model.build_vocab(wiki_vectors)
#using model to get wordmover's distance
distance_results = pd.DataFrame([[base_model.wv.wmdistance(p1, p2) for p2 in data.Messages.str.split()]
for p1 in data.Messages.str.split()]
, columns = data.ID
, index = data.ID )
However, this takes a very long time to run and won't be feasible on larger datasets. I have tried using modin.pandas but that didn't change anything. I have looked into multiprocessing but I need to get the results in the dataframe.
I thought maybe use swifter and create a function to use apply with but I can't seem to get it right.
def compute_wmd(col):
return [base_model.wv.wmdistance(p1, p2) for p2 in data[col].str.split() for p1 in data[col].str.split()]
data['wmd'] = data['Messages'].swifter.apply(compute_wmd)
Any help would be greatly appreciated!