Word2Vec: change of parameter, same results

Viewed 89

I'm trying to train Word2Vec models and I would like to create an embedding to average the results over different models. The problem is that I am not getting the results I expected. In fact, even if I change the parameters I end up I am getting odd results.

The corpus is developed in this way, where 'documents' is a list of tweets.

word_corpus = [[str(token).lower() +'_' + str(token.pos_) for token in nlp(sentence) if token.pos_ in ('NOUN', 'VERB', 'ADJ') and len(str(token))>1] for sentence in documents]

Here the two models:

# initialize model
w2v_model1 = Word2Vec(vector_size=100, # vector size 
                     window=3, # window for sampling
                     sample=0.01, # subsampling rate
                     epochs=10, # iterations
                     negative=10, # negative samples
                     min_count=11, # minimum threshold
                     workers=-1, # parallelize to all cores
                     hs=0 # no hierarchical softmax
)

# build the vocabulary
w2v_model1.build_vocab(corpus)

# train the model
w2v_model1.train(corpus, total_examples=w2v_model1.corpus_count, epochs=w2v_model1.epochs)


#####


w2v_model2 = Word2Vec(vector_size=100, # vector size 
                     window=7, # window for sampling
                     sample=0.01, # subsampling rate
                     epochs=120, # iterations
                     negative=3, # negative samples
                     min_count=100, # minimum threshold
                     workers=-1, # parallelize to all cores
                     hs=0 # no hierarchical softmax
)

# build the vocabulary
w2v_model2.build_vocab(corpus)

# train the model
w2v_model2.train(corpus, total_examples=w2v_model2.corpus_count, epochs=w2v_model2.epochs)

and to evaluate them i am using the following syntax:

emb_df1 = (pd.DataFrame([w2v_model1.wv.get_vector(str(n)) for n in w2v_model1.wv.key_to_index],index = w2v_model1.wv.key_to_index)).T
emb_df2 = (pd.DataFrame([w2v_model2.wv.get_vector(str(n)) for n in w2v_model2.wv.key_to_index],index = w2v_model2.wv.key_to_index)).T

And I get this results. Results from the two different Word2Vec

As you can see the number of words gets different but what seems odd to me is that, for words that are analyzed in both models, I get exactly the same coordinates and I can't understand why. If i well understood its working mechanism the results are provided after some randomization steps and so it should be basically impossible to get the same results everytime, so I am not able to understand to what could be due this error.

0 Answers
Related