I'm trying to understand how to classify a document based on named entities found by earlier pipeline components rather than just the raw text.
Say I have the document "Gross Pay $50. Net Pay $40. Tax $10"
I want to classify the whole text as a PAYSLIP in a multilabel textcat.
With some custom EntityRuler patterns I can easily predict the document entity labels as something like: "Gross Pay [MONEY]. Net Pay [MONEY]. Tax [MONEY]"
My question is, how do I use these labels (stored in doc.ents / Token.ent_type) as features to train a TextCategorizer so it only cares whether a token is MONEY and doesn't distinguish between the different quantities ($50, $40, $10) when predicting a category? ie, how do I classify documents based on token.ent_type and not token.text for all or some of the documents' tokens?
I'm using spaCy 3.2