How can I modify the Spacy English tokenizer so that it will split on, and split apart, specific pairs of punctuation:
import spacy
nlp = spacy.load('en_core_web_md')
doc = nlp("running.(together")
# desired outcome
assert( [t.text for t in doc] == ["running", ".", "(", "together"])
what is currently get, is just one token, "running.(together".
By modify, I mean: do all of the current English tokenization, but also split this run-on case too.