Is there a way to map SpaCy NER labels to new values and evaluate classification performance?

Viewed 275

I would like to map the outputs of a SpaCy NER model to new values.

For instance, SpaCy may assign the label 'LOC' or 'GPE' to a named entity, both referring to something geographical. For my use case, I have a corpus of documents containing named entities assigned the tag 'GEOGRAPHICAL'. I would like to evaluate a SpaCy NER model against this data set. Hence, I would like to map 'LOC' and 'GPE' to 'GEOGRAPHICAL' and be able to evaluate the performance.

Here is some code I am using to try and get this set up:

import spacy
from spacy.scorer import Scorer
from spacy.tokens import Doc
from spacy.training.example import Example
import en_core_web_trf

nlp = en_core_web_trf.load()

data = [
    ('Who is Shaka Khan and United Kingdom?',
     {'entities': [(7, 17, 'PERSON'), (22, 36, 'GPE')]}),
    ('I like London and Berlin and Mount Everest.',
     {'entities': [(7, 13, 'GPE'), (18, 24, 'GPE'), (29, 42, 'LOC')]})
]

examples = []
scorer = Scorer()
for text, annotations in data:
    doc = nlp.make_doc(text)
    example = Example.from_dict(doc, annotations)
    example.predicted = nlp(example.predicted)
    examples.append(example)

scorer.score(examples)

This gives the output:

{'ents_f': 1.0,
 'ents_p': 1.0,
 'ents_per_type': {'GPE': {'f': 1.0, 'p': 1.0, 'r': 1.0},
 'LOC': {'f': 1.0, 'p': 1.0, 'r': 1.0},
 'PERSON': {'f': 1.0, 'p': 1.0, 'r': 1.0}},
 'ents_r': 1.0,
 'token_acc': 1.0,
 'token_f': 1.0,
 'token_p': 1.0,
 'token_r': 1.0}

But I would like to have this instead:

data = [
    ('Who is Shaka Khan and United Kingdom?',
     {'entities': [(7, 17, 'PERSON'), (22, 36, 'GEOGRAPHICAL')]}),
    ('I like London and Berlin and Mount Everest.',
     {'entities': [(7, 13, 'GEOGRAPHICAL'), (18, 24, 'GEOGRAPHICAL'), (29, 42, 'GEOGRAPHICAL')]})
]

Followed by this output:

{'ents_f': 1.0,
 'ents_p': 1.0,
 'ents_per_type': {'GEOGRAPHICAL': {'f': 1.0, 'p': 1.0, 'r': 1.0},
 'PERSON': {'f': 1.0, 'p': 1.0, 'r': 1.0}},
 'ents_r': 1.0,
 'token_acc': 1.0,
 'token_f': 1.0,
 'token_p': 1.0,
 'token_r': 1.0}

Can this be achieved? How can this be achieved?


Update: My idea at the moment is to take the output of nlp(example.predicted) and map the labels like this:

    data = [
        ('Who is Shaka Khan and United Kingdom?',
         {'entities': [(7, 17, 'PERSON'), (22, 36, 'GEOGRAPHICAL')]}),
        ('I like London and Berlin and Mount Everest.',
         {'entities': [(7, 13, 'GEOGRAPHICAL'), (18, 24, 'GEOGRAPHICAL'), (29, 42, 'GEOGRAPHICAL')]})
    ]

    mappings = {'LOC':'GEOGRAPHICAL', 'GPE':'GEOGRAPHICAL', 'PERSON':'PERSON'}
    
    examples = []
    scorer = Scorer()
    for text, annotations in data:
        doc = nlp.make_doc(text)
        example = Example.from_dict(doc, annotations)
        example.predicted = nlp(example.predicted)
    
        ents = list(example.predicted.ents)
    
        for i, ent in enumerate(ents):
            ents[i] = Span(example.predicted, ent.start, ent.end, label=mappings[ent.label_])
    
        example.predicted.ents = ents
    
        examples.append(example)
    
    scorer.score(examples)

This seems to be producing the correct output. Maybe the focus of this question is now determining if there is a more elegant way to do this.

0 Answers
Related