word not in vocabulary

Viewed 720

First time using word2vec and the file I am working with is in XML format. I want to iterate through the patents to find each Title then apply word2vec to see if there are similar words(to indicate similar titles).

So far I have parsed the XML file using Element tree to retrieve each title, then I have applied sent_tokenizer followed by tweet tokenizer to return a list of sentences where each word has been tokenized (not sure if this was the best method). I then put the tokenized sentenses into my word2vec model and tested with one word to see if it returned a vector. This seems to only work for a word in the first sentence. I'm not sure it is recognising all the sentences?

    import numpy as np
    import pandas as pd
    import gensim
    import nltk
    import xml.etree.ElementTree as ET
    from gensim.models.word2vec import Word2Vec
    from nltk.tokenize import word_tokenize
    from nltk.tokenize import sent_tokenize
    from nltk.corpus import stopwords
    from nltk.tokenize import TweetTokenizer, sent_tokenize

    tree = ET.parse('6785.xml')
    root = tree.getroot()

    for child in root.iter("Title"):
        Patent_Title = child.text
        sentence = Patent_Title
        stopWords = set(stopwords.words('english'))
        tokens = nltk.sent_tokenize(sentence)
        print(tokens)

        tokenizer_words = TweetTokenizer()
        tokens_sentences = [tokenizer_words.tokenize(t) for t in tokens]
        #print(tokens_sentences)

        model = gensim.models.Word2Vec(tokens_sentences, min_count=1,size=32)
        words = list(model.wv.vocab)
        print(words)
        print(model['Solar'])

I would expect it to identify the word 'solar' in a sentence and print out the vector then I could look for similar words. I am receiving the error:

word 'Solar' not in vocabulary"

2 Answers

Just handle the errors as exceptions on first loop occurence.

# print(model['Solar'])
try:
    print(model['Solar'])
except Exception as e:
    pass

Working code :

import numpy as np
import pandas as pd
import gensim
import nltk
import xml.etree.ElementTree as ET
from gensim.models.word2vec import Word2Vec
from nltk.tokenize import word_tokenize
from nltk.tokenize import sent_tokenize
from nltk.corpus import stopwords
from nltk.tokenize import TweetTokenizer, sent_tokenize

tree = ET.parse('6785.xml')
root = tree.getroot()

for child in root.iter("Title"):
    Patent_Title = child.text
    sentence = Patent_Title
    stopWords = set(stopwords.words('english'))
    tokens = nltk.sent_tokenize(sentence)
    print(tokens)

    tokenizer_words = TweetTokenizer()
    tokens_sentences = [tokenizer_words.tokenize(t) for t in tokens]
    #print(tokens_sentences)

    model = gensim.models.Word2Vec(tokens_sentences, min_count=1,size=32)
    words = list(model.wv.vocab)
    print(words)
    try:
        print(model['Solar'])
    except Exception as e:
        pass

It is simply because Solar is not in your corpus.

Word2Vec tries to generate word vectors for each word in your tokens_sentences. If the training corpus didn't include the word/token that you try to look up, word2vec would not have the word vector for that word and that is why you got an error.

Advice: try to make your text data case-insensitive. That is, make all the text lower case (upper case works too but not the convention.)

Related