Difference between spacy v3 en_core_web_trf pipeline and en_core_web_lg pipeline

Viewed 1605

I am doing some performance tests with spacy version 3 for right sizing my instances in production. I am observing the following

Observation:

Model name Time without NER Time with NER Comments
en_core_web_lg 4.89 seconds 21.9 seconds NER adds 350% to the original time
en_core_web_trf 43.64 seconds 52.83 seconds NER adds just 20% to the original time

Why is there no significant difference between the with NER and without NER scenarios in the case of the transformer model? Is NER just an incremental task after POS tagging in the case of en_core_web_trf?

Test environment: GPU instance

Test code:

import spacy

assert(spacy.__version__ == '3.0.3')
spacy.require_gpu()
texts = load_sample_texts()  # loads 10,000 texts from a file
assert(len(texts) == 10000)

def get_execution_time(nlp, texts, N):
    return timeit.timeit(stmt="[nlp(text) for text in texts]", 
                           globals={'nlp': nlp, 'texts': texts}, number=N) / N


#  load models
nlp_lg_pos = spacy.load('en_core_web_lg', disable=['ner', 'parser'])
nlp_lg_all = spacy.load('en_core_web_lg')
nlp_trf_pos = spacy.load('en_core_web_trf', disable=['ner', 'parser'])
nlp_trf_all = spacy.load('en_core_web_trf')

#  get execution time
print(f'nlp_lg_pos = {get_execution_time(nlp_lg_pos, texts, N=1)}')
print(f'nlp_lg_all = {get_execution_time(nlp_lg_all, texts, N=1)}')
print(f'nlp_trf_pos = {get_execution_time(nlp_trf_pos, texts, N=1)}')
print(f'nlp_trf_all = {get_execution_time(nlp_trf_all, texts, N=1)}')
0 Answers
Related