How to train a TextCategorizer using the entities matched by NER or EntityRuler in spaCy?

Viewed 26

I'm trying to understand how to classify a document based on named entities found by earlier pipeline components rather than just the raw text.

Say I have the document "Gross Pay $50. Net Pay $40. Tax $10"

I want to classify the whole text as a PAYSLIP in a multilabel textcat.

With some custom EntityRuler patterns I can easily predict the document entity labels as something like: "Gross Pay [MONEY]. Net Pay [MONEY]. Tax [MONEY]"

My question is, how do I use these labels (stored in doc.ents / Token.ent_type) as features to train a TextCategorizer so it only cares whether a token is MONEY and doesn't distinguish between the different quantities ($50, $40, $10) when predicting a category? ie, how do I classify documents based on token.ent_type and not token.text for all or some of the documents' tokens?

I'm using spaCy 3.2

0 Answers
Related