I have used google's implementation of electra to train a model from scratch. For the pretraining I have followed this tutorial -- with modifications of course, since it uses google's implementation for BERT. After having finished the training & also having converted the electra discriminator checkpoint to huggingface format using this script, I am trying to load the model in order to get the embeddings for some sentences.
First, I create the ElectraConfig, and then I create a tokenizer using the vocab.txt I have created (based on the aforementioned tutorial)
config = ElectraConfig(vocab_size=100000,
embedding_size=768,
hidden_size=768,
num_hidden_layers=12,
num_attention_heads=4,
intermediate_size=3072,
hidden_act="gelu",
hidden_dropout_prob=0.1,
attention_probs_dropout_prob=0.1,
max_position_embeddings=512,
position_embedding_type="absolute")
tokenizer = ElectraTokenizer.from_pretrained(path_to_vocab)
My problem is that the tokenizer encodes all tokens in all sentences as unknown. Do I need to convert the vocab.txt file? Is the creation of the vocab.txt file described in the tutorial incompatible with huggingface?