I am attempting to fine-tune BERT in tensorflow following this official guide with the goal of feeding the output further into LSTM/GRU. I am able to run the fine-tuning but the output shapes I am getting from the bert_encoder are [num_samples, hidden_units] and [num_samples, 1, 768]. I believe these are pooled and sequence outputs respectively but I am confused why the sequence output is not [num_samples, max_seq_length, hidden_units].
Running this code after replacing bert_classifier with bert_encoder on compile and fit:
bert_encoder([glue_train["input_word_ids"][0:10],
glue_train["input_mask"][0:10],
glue_train["input_type_ids"][0:10]])
produces:
[<tf.Tensor: shape=(10, 1, 768), dtype=float32, numpy= ...>, <tf.Tensor: shape=(10, 768), dtype=float32, numpy= ...>]
Since I am passing to sequence models, I need to get the sequence output but I keep getting only 1 shape for the sequence length. I've been trying to understand why but can't find anything on it. Any help and clarifications would be appreciated. Thanks!