I'm trying to find entities on websites text. So I've copied 20 sites into a text file (all in one line) and tagged entities manually. (according to this tutorial: https://www.machinelearningplus.com/nlp/training-custom-ner-model-in-spacy/)
Usually, the text files contain 5000+ characters and I've tagged two entities per each file.
I've got Spacy 3.2.1 and I'm using nlp = spacy.load("en_core_web_sm").
The print of the losses:
Losses {'ner': 438.16809472077654}
Losses {'ner': 448.97231569240785}
Losses {'ner': 470.66727516808453}
Losses {'ner': 477.0379003697807}
Losses {'ner': 8.354419636216686}
Losses {'ner': 86.56611033267922}
...
Losses {'ner': 0.026654468748418734}
Losses {'ner': 0.026672022747312552}
Losses {'ner': 0.026672030436685996}
Losses {'ner': 0.027033395785247216}
Losses {'ner': 0.027033751640550253}
Losses {'ner': 0.02703389312135557}
Losses {'ner': 0.029515604829788607}
When I use this model to find an entity on another text, it finds nothing. (even though I can see the entities in the text).
doc = nlp(text)
print("Entities", [(ent.text, ent.label_) for ent in doc.ents])
displacy.serve(doc, style="ent")
This code just displays:
Entities []
/home/python3.8/site-packages/spacy/displacy/__init__.py:200: UserWarning: [W006] No entities to visualize found in Doc object. If this is surprising to you, make sure the Doc was processed using a model that supports named entity recognition, and check the `doc.ents` property manually if necessary.
warnings.warn(Warnings.W006)
Using the 'ent' visualizer
Serving on http://0.0.0.0:5000
All the examples that I found online use a different text to entity ratio: 1 short sentence with 1 - 2 entities inside. I'm using 1 long text with 1 - 2 entities inside.
Could that be the issue? And if so, what do I need to do? Prepare more training data?