I am using the text classification pipeline from transformers like this.
from transformers import pipeline
model_root = "<Path to my model>"
classifier = pipeline("text-classification", model=model_root, tokenizer=model_root)
results = classifier(["list of sentences"])
My problem is that the sentences which reach this part of the code are split into words. So basically I have a lists of tokens nested into another list like this.
[["this", "is", "sentence", "one"],
["this", "is", "sentence", "two"],
["this", "is", "sentence", "three"]]
Is there a way to tell the pipeline to do the tokenization with is_split_into_words set to True? I want to avoid joining the lists into sentences.