Huggingface tokenizer always encodes as unknown - Conversion of vocab.txt?

Viewed 93

I have used google's implementation of electra to train a model from scratch. For the pretraining I have followed this tutorial -- with modifications of course, since it uses google's implementation for BERT. After having finished the training & also having converted the electra discriminator checkpoint to huggingface format using this script, I am trying to load the model in order to get the embeddings for some sentences.

First, I create the ElectraConfig, and then I create a tokenizer using the vocab.txt I have created (based on the aforementioned tutorial)

config = ElectraConfig(vocab_size=100000,
                       embedding_size=768,
                       hidden_size=768,
                       num_hidden_layers=12,
                       num_attention_heads=4,
                       intermediate_size=3072,
                       hidden_act="gelu",
                       hidden_dropout_prob=0.1,
                       attention_probs_dropout_prob=0.1,
                       max_position_embeddings=512,
                       position_embedding_type="absolute")

tokenizer = ElectraTokenizer.from_pretrained(path_to_vocab)

My problem is that the tokenizer encodes all tokens in all sentences as unknown. Do I need to convert the vocab.txt file? Is the creation of the vocab.txt file described in the tutorial incompatible with huggingface?

0 Answers
Related