Mixed Unigram Bigram word2vec embedding

Viewed 930

I am trying to build an embedding for a corpus using Python's gensim's word2vec implementation. The catch is that I wish to have in the same embedding all of the unigrams and bigrams for the corpus. Is there a way to embed in the same space both unigrams and bigrams?

1 Answers

You can do it with Phrases model in gensim

from gensim.models.phrases import Phrases, Phraser

#documents is list is list of tokens from your text
bigram  = Phrases(documents, min_count=2)
trigram   = Phrases(bigram[documents], min_count=1)
Related