I am working in Python with a dictionary in this form:
{ 1:[(word1, word2), (word3, word4)], 2:[(word5, word6), (word7, word8), (word9, word10)], 3:[(word11, word12), (word13, word14)] }
I am computing the cosine similarity of each word pair (using Word2Vec), indexed by key as following:
> def get_sim(data, key=int):
> for key in data:
> for w1, w2 in temp[1]:
> print(key, w1, w2, wv.similarity(w1, w2))
get_sim(temp)
from which I get this kind of result:
1, word1, word2, cos_sim
My question: for all similarity scores associated to each key (cos_sim values), I want to compute the mean value (like a final score). Could list comprehension help?
My practical question: which would be a suitable module to work with the datatype above? I tried with json but it transforms my keys (and cos_sim values) in strings, which I definitely don't want. Even using int(key) doesn't help. Pandas and numpy, on the other hand, don't work very good with strings which are my 2nd and 3rd entry in the line.
I am relatively new to programming, so thank you very much for any hint which will me help further!