Argument "never_split" not working on bert tokenizer

Viewed 945

I used the never_split option and tried to retain some tokens. But the tokenizer still divide them into wordpieces.

tokenizer = BertTokenizer.from_pretrained('bert-base-uncased', never_split=['lol'])
tokenizer.tokenize("lol That's funny")
['lo', '##l', 'that', "'", 's', 'funny']

Do I miss anything here?

1 Answers

I would call this a bug or at least not good documented. The never_split argument is only considered when you use the BasicTokenizer (which is part of the BertTokenizer).

You are calling the tokenize function from your specific model (bert-base-uncased) and this considers only his vocabulary (as I would expect). In order to prevent splitting of certain tokens, they must be part of the vocabulary (you can extend the vocabulary with the method add_tokens).

I think the example below shows what I am trying to say:

from transformers import BertTokenizer

text = "lol That's funny lool"

tokenizer = BertTokenizer.from_pretrained('bert-base-uncased', never_split=['lol'])
#what you are doing
print(tokenizer.tokenize(text))

#how it is currently working
print(tokenizer.basic_tokenizer.tokenize(text))

#how you should do it
tokenizer.add_tokens('lol')
print(tokenizer.tokenize(text))

Output:

['lo', '##l', 'that', "'", 's', 'funny', 'lo', '##ol']
['lol', 'that', "'", 's', 'funny', 'lool']
['lol', 'that', "'", 's', 'funny', 'lo', '##ol']
Related