i need to index some document with custom tokenizer. my sample doc is look like this:
"I love to live in New York"
and list of expressions is:
["new york", "good bye", "cold war"]
is there any way to tokenize string normally but do not tokenize my dataset?
["I", "love", "to", "live", "in", "New York"]