I have setup a Returnn Transformer Model for NMT, which I want to train with an additional loss for every encoder/decoder attention head h on every decoder layer l (in addition to the vanilla Cross Entropy loss), i.e.:
loss = CrossEntropyLoss + sum_{Layer l=1,...,6} sum_{Head h=1,...,8} (lambda * AttentionLoss(l, h))
for some scalar lambda. I implemented the attention loss itself as a eval-Layer, using the loss=as_is option, that returns a single number for each batch (that is the value of lambda * AttentionLoss(l, h).
As a test, I also implemented a version where I have one loss for each layer l, equivalent to lambda * sum_{Head h=1,...,8} AttentionLoss(l, h) to reduce the number of losses as I noticed a decrease in performance and as log files were getting very large as Returnn prints every loss for each batch.
However, I got very different results for both implementations: A model trained with one loss per layer AND head performs consistently better. I tried this with multiple training runs.
To investigate this, I tried a training run where I set the parameter lambda=0.0, i.e. effectively disabled the attention loss. And even here, in comparison to the baseline without any additional losses, a model trained with these additional 6 losses all outputting a constant 0 performs noticably worse, see this table:
+--------------------------------------------+-------------+-------------+
| | Dev Set | Test Set |
+--------------------------------------------+------+------+------+------+
| | BLEU | TER | BLEU | TER |
+--------------------------------------------+------+------+------+------+
| Only Cross Entropy Loss | 35.7 | 51.4 | 34.2 | 53.5 |
+--------------------------------------------+------+------+------+------+
| + One loss per layer and head (lambda 0) | 35.5 | 51.5 | 33.9 | 53.7 |
+--------------------------------------------+------+------+------+------+
| + One loss per layer (lambda 0) | 35.4 | 51.8 | 33.5 | 54.2 |
+--------------------------------------------+------+------+------+------+
| + Simplified One loss per layer (lambda 0) | 35.1 | 52.0 | 33.5 | 54.3 |
+--------------------------------------------+------+------+------+------+
Here, the "simplified" version is implemented exactly like this:
'dec_01_weight_loss': {
'class': 'eval', 'eval': '0.0 * tf.reduce_sum(source(0, auto_convert=False))',
'from': ['dec_01_att_weights'], 'loss': 'as_is',
'out_type': { 'batch_dim_axis': None, 'dim': None, 'dtype': 'float32', 'feature_dim_axis': None,
'shape': (), 'time_dim_axis': None}}
while the actual loss I use is a bit more complicated, I uploaded my full config files here. (Here the loss layer is called dec_01_att_weight_variance etc.)
And all lambda=0.0 implementions mentioned above output the value 0.0 for all additional losses in every training step:
train epoch 1, step 0, cost:output/dec_01_weight_loss 0.0, cost:output/dec_02_weight_loss 0.0, cost:output/dec_03_weight_loss 0.0, [....], cost:output/output_prob 8.541749455164052, error:decision 0.0, error:output/output_prob 0.9999999680730979, loss 8.5417 49, max_mem_usage:GPU:0 1.2GB, mem_usage:GPU:0 1.2GB, 3.999 sec/step, elapsed 0:00:38, exp. remaining 1:30:00, complete 0.71%
What is going on here? Is there any explanation why the models behave differently, why does an additional loss with a constant value 0.0 change the model behavior?
I am using TF 1.15.0 (v1.15.0-0-g590d6eef7e), Returnn 20200613.152716--git-23332ca, using Python 3.8.0 with CUDA 10.1.
Followup Update: I tested the same config using pre-training, where I would disable my loss completely for the first n-1 (here e.g. n=50) checkpoints using the following code:
def custom_construction_algo(idx, net_dict):
if idx == 0:
for lay in range(1, 7):
del net_dict["output"]["unit"]["dec_%02i_att_loss" % lay]
return net_dict
else:
return None
pretrain = {"repetitions": 49, "construction_algo": custom_construction_algo}
In the log file, for the first n-1 checkpoints I (correctly) only see the CE loss being reported.
Here I am showing my Dev BLEU at the last checkpoint trained without the additional loss (i.e. n-1, here 49), each experiment run multiple times:
- Baseline (no additional loss): 31.8, 31.7, 31.7 BLEU
- One loss per layer disabled with pretraining: 29.2, 29.0, 28.5 BLEU
- One loss per layer with
lambda=0.0(as in original question): 28.8, 28.7 BLEU - One loss per layer AND head with
lambda=0.0(as in original question): 31.8 BLEU
From my understanding, the TF graph for pre-training config and the baseline should be identical up to checkpoint n=50. Yet, they perform very differently. What is going on?
The full config I used for this kind of pre-training can be found here. The heads of the corresponding log files are found here. I am using NewbobMultiEpoch with Adam:
learning rate control: NewbobMultiEpoch(num_epochs=9, update_interval=1, relative_error_threshold=0, learning_rate_decay_factor=0.7, learning_rate_growth_factor=1.0), epoch data: , error key: None
Create optimizer <class 'tensorflow.python.training.adam.AdamOptimizer'> with options {'beta1': 0.9, 'beta2': 0.999, 'epsilon': 1e-08, 'learning_rate': <tf.Variable 'learning_rate:0' shape=() dtype=float32_ref>}.
For all reported experiments, the learning rate does not decrease until checkpoints larger than 100, staying constant at initial 10^-4.
EDIT: I did a mistake and accidentally used a different Returnn version across my experiments. The Returnn I used for my experiments with additional losses seems to have contained some local changes I made. When rerunning a baseline with the new version, it performed significantly worse - very similar to the other BLEU values documented here. A subtle bug in one of my Returnn versions - thats all that was to this issue.