How can I pass is_split_into_words option to a huggingface transformers pipeline?

Viewed 19

I am using the text classification pipeline from transformers like this.

from transformers import pipeline

model_root = "<Path to my model>"

classifier = pipeline("text-classification", model=model_root, tokenizer=model_root)

results = classifier(["list of sentences"])

My problem is that the sentences which reach this part of the code are split into words. So basically I have a lists of tokens nested into another list like this.

[["this", "is", "sentence", "one"],
 ["this", "is", "sentence", "two"],
 ["this", "is", "sentence", "three"]]

Is there a way to tell the pipeline to do the tokenization with is_split_into_words set to True? I want to avoid joining the lists into sentences.

0 Answers
Related