I'm trying to train Word2Vec models and I would like to create an embedding to average the results over different models. The problem is that I am not getting the results I expected. In fact, even if I change the parameters I end up I am getting odd results.
The corpus is developed in this way, where 'documents' is a list of tweets.
word_corpus = [[str(token).lower() +'_' + str(token.pos_) for token in nlp(sentence) if token.pos_ in ('NOUN', 'VERB', 'ADJ') and len(str(token))>1] for sentence in documents]
Here the two models:
# initialize model
w2v_model1 = Word2Vec(vector_size=100, # vector size
window=3, # window for sampling
sample=0.01, # subsampling rate
epochs=10, # iterations
negative=10, # negative samples
min_count=11, # minimum threshold
workers=-1, # parallelize to all cores
hs=0 # no hierarchical softmax
)
# build the vocabulary
w2v_model1.build_vocab(corpus)
# train the model
w2v_model1.train(corpus, total_examples=w2v_model1.corpus_count, epochs=w2v_model1.epochs)
#####
w2v_model2 = Word2Vec(vector_size=100, # vector size
window=7, # window for sampling
sample=0.01, # subsampling rate
epochs=120, # iterations
negative=3, # negative samples
min_count=100, # minimum threshold
workers=-1, # parallelize to all cores
hs=0 # no hierarchical softmax
)
# build the vocabulary
w2v_model2.build_vocab(corpus)
# train the model
w2v_model2.train(corpus, total_examples=w2v_model2.corpus_count, epochs=w2v_model2.epochs)
and to evaluate them i am using the following syntax:
emb_df1 = (pd.DataFrame([w2v_model1.wv.get_vector(str(n)) for n in w2v_model1.wv.key_to_index],index = w2v_model1.wv.key_to_index)).T
emb_df2 = (pd.DataFrame([w2v_model2.wv.get_vector(str(n)) for n in w2v_model2.wv.key_to_index],index = w2v_model2.wv.key_to_index)).T
As you can see the number of words gets different but what seems odd to me is that, for words that are analyzed in both models, I get exactly the same coordinates and I can't understand why. If i well understood its working mechanism the results are provided after some randomization steps and so it should be basically impossible to get the same results everytime, so I am not able to understand to what could be due this error.
