PyTorch tokenizers: how to truncate tokens from left?

Viewed 297

As we can see in the below code snippet, specifying max_length and truncation for a tokenizer cuts excess tokens from the left:

tokenizer("hello, my name", truncation=True, max_length=6).input_ids

> [0, 42891, 6, 127, 766, 2]

tokenizer("hello, my name", truncation=True, max_length=4).input_ids

> [0, 42891, 6, 2]

The different tokenization strategies like only_second, only_first, longest_first, all seem to cut from the right. Is there a way to cut from the left? So that the tokens would be [0, 127, 766, 2] in the second example?

0 Answers
Related