Any reason to save a pretrained BERT tokenizer?

Viewed 1770

Say I am using tokenizer = BertTokenizer.from_pretrained('bert-base-uncased', do_lower_case=True), and all I am doing with that tokenizer during fine-tuning of a new model is the standard tokenizer.encode()

I have seen in most places that people save that tokenizer at the same time that they save their model, but I am unclear on why it's necessary to save since it seems like an out-of-the-box tokenizer that does not get modified in any way during training.

3 Answers

In your case, if you are using tokenizer only to tokenize the text (encode()), then you need not have to save the tokenizer. You can always load the tokenizer of the pretrained model.

However, sometimes you may want to use the tokenizer of the pretrained model, then add new tokens to it's vocabulary, or redefine the special symbols such as '[CLS]', '[MASK]', '[SEP]', '[PAD]' or any such special tokens. In this case, since you have made the changes to the tokenizer, it will be useful to save the tokenizer for the future use.

You can always wake up the tokenizer with:

tokenizer = BertTokenizer.from_pretrained('bert-base-uncased', do_lower_case=True)

This may be just part of the routine, that is not so needed.

Tokenizers create their vocabulary based on the frequency of words (or subwords as in byte pair encoding) in a training corpus. The same tokenizer may have a different vocabulary depending on the corpus on which it is trained.

For this reason you probably want to save the tokenizer after "training" it on a corpus and subsequently training a model that used that tokenizer.

The Huggingface Tokenizer Summary covers how these vocabularies are built up.

Related