How not to get "datum" as the lemma for "data" when using Spacy?

Viewed 639

I've run into a quite common word "data" which gets assigned a lemma "datum" from lookups exceptions table spacy uses. I understand that the lemma is technically correct, but in today's english, "data" in its basic form is just "data". I am using the lemmas to get a sort of keywords from text and if I have a text about data, I can't possibly tag it with "datum". I was wondering if there is another way to arrive at plain "data" then constructing another "my_exceptions" dictionary used for overriding post-processing. Thanks for any suggestions.

2 Answers

You could use Lemminflect which works as an add-in pipeline component for SpaCy. It should give you better results.

To use it with SpaCy, just import lemminflect and call the new ._.lemma() function on the Token, ie.. token._.lemma(). Here's an example..

import lemminflect
import spacy
nlp = spacy.load('en_core_web_sm')
doc = nlp('I got the data')
for token in doc:
    print('%-6s %-6s %s' % (token.text, token.lemma_, token._.lemma()))

I      -PRON- I
got    get    get
the    the    the
data   datum  data

Lemminflect has a prioritized list of lemmas, based on occurrence in corpus data. You can see all lemmas with...

print(lemminflect.getAllLemmas('data'))

{'NOUN': ('data', 'datum')}

It's relatively easy to customize the lemmatizer once you know where to look. The original lemmatizer tables are from the package spacy-lookups-data and are loaded in the model under nlp.vocab.lookups. You can use a local install of spacy-lookups-data to customize the tables for new/blank models, but if you just want to make a few modifications to the entries for an existing model, you can modify the lemmatizer tables on the fly.

Depending on whether your pipeline includes a tagger, the lemmatizer may be referring to rules+exceptions (with a tagger) or to a simple lookup table (without a tagger), both of which include an exception that lemmatizes data to datum by default. If you remove this exception, you should get data as the lemma for data.

For a pipeline that includes a tagger (rule-based lemmatizer)

# tested with spaCy v2.2.4
import spacy

nlp = spacy.load("en_core_web_sm")

# remove exception from rule-based exceptions
lemma_exc = nlp.vocab.lookups.get_table("lemma_exc")
del lemma_exc[nlp.vocab.strings["noun"]]["data"]

assert nlp.vocab.morphology.lemmatizer("data", "NOUN") == ["data"]

# "data" with the POS "NOUN" has the lemma "data"
doc = nlp("data")
doc[0].pos_ = "NOUN" # just to make sure the POS is correct
assert doc[0].lemma_ == "data"

For a pipeline without a tagger (simple lookup lemmatizer)

import spacy

nlp = spacy.blank("en")

# remove exception from lookups
lemma_lookup = nlp.vocab.lookups.get_table("lemma_lookup")
del lemma_lookup[nlp.vocab.strings["data"]]

assert nlp.vocab.morphology.lemmatizer("data", "") == ["data"]

doc = nlp("data")
assert doc[0].lemma_ == "data"

For both: save model for future use with these modifications included in the lemmatizer tables

nlp.to_disk("/path/to/model")

Also be aware that the lemmatizer uses a cache, so make any changes before running your model on any texts or you may run into problems where it returns lemmas from the cache rather than the updated tables.

Related