I think this might could help finetune a pretrained model using tensorflow_hub with MLM task, we can use keras_nlp from official.nlp.
Here is a simple example:
bert_encoder_scratch = get_transformer_encoder(config)
bert_encoder_pretrain = hub.KerasLayer(hub.load(module_url), trainable=True)
if you are training from scratch:
mlm_model = keras_nlp.layers.MaskedLM(embedding_table=bert_encoder_scratch.get_embedding_table())
So in this way, what lacks for a hub model is to find the embedding table. Find the variable word_embedding by:
bert_encoder_pretrain.trainable_variables
or you can use tf.get_variable(). Here I simply write bert_encoder_pretrain.trainable_variables[0]
The mlm task would be:
mlm_model = keras_nlp.layers.MaskedLM(embedding_table=bert_encoder_pretrain.trainable_variables[0])
Then you get your seq_output from the model, and also masked_positions, here I define it as:
masked_positions = tf.broadcast_to(tf.range(0, seq_length), [batch_size, seq_length])
You will get the mlm_logits from the model by:
mlm_logits = mlm_model(seq_output, masked_positions)
I have tested this approach in my model.