Suppose I know ahead of time the character-level sentence boundaries in a document:
text = "The cat chased the mouse. The mouse ran away."
boundaries = [(0, 25), (26, 45)]
for start, end in boundaries:
print(text[start:end])
Is there a way that I can tell Spacy to use these boundaries? From what I can gather in the official docs and elsewhere on SO, the hooks provided seem more suited to support custom stateless rules that apply at the word (token) level.