With the Transformer model, especially with the BERT, does it make sense (and would it be statistically correct) to programmatically forbid the model to result with the special tokens as predictions? How is that in the original implementations? During convergence the models have to learn not to predict these but would this intervention help (or the opposite)?
- I would consider the [MASK], [CLS] tokens mainly
- [PAD] token could have some sense as well (but that not in all situations)