Can NLTK pos tagger recognize contractions correctly?

Viewed 76

I want to know if I need to write a de-contraction function before sending a given text to NLTK's pos tagger. I am reluctant to tokenize words because they might end up like (don't='do',"'nt") which I suspect would make pos tagging more difficult.

In short, my questions are: Does nltk's pos tagger recognize most contractions (from my limited experience it seems to work well w/o word tokenization)? Will word tokenization (as opposed to simple word splitting) improve or impair the process? Would it just be easier for me to write a de-contraction function? Are there any other pos taggers that recognize contractions?

example_text="I can't and I won't go to the park because I don't like grass."

2 Answers

I'm currently running into similar problems, and as much I'd like to get the hang of NLTK (for synsets more than anything), I find Spacy to be more user-friendly for a relative newbie like me, and more robust when handling contractions.

Installation:

pip install spacy

python -m spacy download en_core_web_sm (or en_core_web_md or en_core_web_lg or en_core_web_trf)

Example:

import spacy
nlp = spacy.load('en_core_web_md')

sent = "I didn't believe it"
tokens_pos = nlp(sent)
for token in tokens_pos:
    print(token.lemma_ + ' ' + token.pos_)

Output:

I PRON
do AUX
not PART
believe VERB
it PRON

As you can see, it lemmatizes and tags at the same time, which need to be done separately with NLTK.

Spacy even works with typically misspeled/lazy forms of the contractions such as cant or dont without the apostrophe. Use regex to make sure your sentences are clean before processing.

Of course nothing is ever 100% perfect. In my output, I actually noticed that I had some rogue tokens tagged as SPACE, which I think might have been a result of double spaces in the original text. It seems to split the string into tokens at each space, treating an adjacent space as a part of speech in its own right, so you might want to add some functionality to filter these out.

Something else I noticed in my experimenting (and I might actually make a post about this), is the way apostrophes are handled.

In my original texts, some contractions had a 6- or 9-shaped apostrophe ā€™ā€˜ typically found at the start and end of quotations, rather than a standard vertical ', and was handled differently by Spacy, so you might also want to make sure these get replaced before you do any NLPing.

I want to know if I need to write a de-contraction function before sending a given text to NLTK's pos tagger.

You do not. The default nltk tagger is trained with text that was tokenized with the default nltk tokenization, and works correctly with text that is tokenized the same way. Anything else would be a bug in the nltk. So if you change the tokenizer you will make performance worse, not better.

If you try your own example you'll see that it correctly tags "ca" and "wo" as MD (modal verb), even though there are no such words in English; I don't particularly like it (why not just tokenize "can't" as "can n't"?), but the tagger certainly knows what to do with it.

>>> nltk.pos_tag(nltk.word_tokenize(example_text))
[('I', 'PRP'), ('ca', 'MD'), ("n't", 'RB'), ('and', 'CC'), ('I', 'PRP'),
 ('wo', 'MD'), ("n't", 'RB'), ('go', 'VB'), ('to', 'TO'), ('the', 'DT'),
 ('park', 'NN'), ('because', 'IN'), ('I', 'PRP'), ('do', 'VBP'), ("n't", 'RB'), 
('like', 'VB'), ('grass', 'NN'), ('.', '.')]

Will the tagger get some things wrong? Definitely. No tagger is perfect. But if you want better performance, you need to find or train a better tagger. You can't "improve" the word tokenizer that the tagger is designed to work with.

PS. You should only pass one (tokenized) sentence at a time to the tagger. If you pass it your entire file as a list of words, you do lose performance unnecessarily. This is how you should do it:

sents = [ nltk.word_tokenize(s) for s in nltk.sent_tokenize(long_text) ]
nltk.pos_tag_sents(sents)
Related