how can i get the number of words that have been influenced by the lemmatization approach in a text?

Viewed 43

For example, in the below sentence where the lemmatizer has affected 5 words, the number 5 should be displayed in the output.

lemmatizer = WordNetLemmatizer()
sentence = "The striped bats are hanging on their feet for best"
print([lemmatizer.lemmatize(w, get_wordnet_pos(w)) for w in nltk.word_tokenize(sentence)])
#> ['The', 'strip', 'bat', 'be', 'hang', 'on', 'their', 'foot', 'for', 'best']
1 Answers

Probably not the most elegant way, but as a workaround you could try to compare every single element of the tokenized sentence and the lemmatized sentence (as far as I know, lemmatization does not remove elements so this should work).

Something like that:

count = 0
for i, el in enumerate(tokenized):
  if el!=lemmatized[i]:
    count+=1

The value of count will be the number of elements that differ from the 2 lists, thus the number of elements affected by lemmatization.

Related