How do I add a custom tokenization rule to spacy for the case of wanting a number and a symbol or word to be tokenized together. E.g. the following sentence:
"I 100% like apples. I like 500g of apples"
is tokenized as follows:
['I', '100', '%', 'like', 'apples', '.', 'I', 'like', '500', 'g', 'of', 'apples']
It would be preferable if it was tokenized like this:
['I', '100%', 'like', 'apples', '.', 'I', 'like', '500g', 'of', 'apples']
The following code was used to generate this:
import spacy
nlp = spacy.load("en_core_web_sm")
text = "I 100% like apples. I like 500g of apples"
print([token.text for token in nlp(text)])