How do loss functions work in Tensorflow? How are they computed from the batches?

Viewed 229

I have had quite a lot of hands-on experience with tensorflow and lately I've been trying to implement a custom loss function. Since I was struggling with it, I tried to implement a simple Mean Absolute Error (MAE) loss. This is my function:

@tf.function
def my_mae(y_true, y_pred):
    return tf.reduce_mean(tf.abs(tf.subtract(y_true, y_pred)), axis=-1)

Compile and fit

Now, this looks pretty accurate to me, but then I compile my model with the following parameters:

model.compile(loss=my_mae,
              optimizer='adam',
              metrics=['mae', 'mse'])

and I start the training with model.fit. The problem is that in the log from the fit function I can see that my MAE and the MAE metric have different values:

Epoch 1/100
483/483 [==============================] - 1s 3ms/step - loss: 3.7004 - mae: 0.7226 - mse: 0.8044 - val_loss: 3.2607 - val_mae: 0.5098 - val_mse: 0.4458
Epoch 2/100
483/483 [==============================] - 1s 3ms/step - loss: 3.4139 - mae: 0.6550 - mse: 0.6687 - val_loss: 3.0994 - val_mae: 0.4907 - val_mse: 0.4207

Am I doing something wrong? Is tensorflow doing something that I don't know?

More experimentation

I also tried to divide by some big number the loss, to see what happened, like in this snippet:

@tf.function
def my_mae(y_true, y_pred):
    return tf.reduce_mean(tf.abs(tf.subtract(y_true, y_pred)), axis=-1) / 1000

but I got the exact same starting loss values:

Epoch 1/100
483/483 [==============================] - 1s 3ms/step - loss: 3.8101 - mae: 0.7587 - mse: 0.8851 - val_loss: 3.3032 - val_mae: 0.5203 - val_mse: 0.4549
Epoch 2/100
483/483 [==============================] - 1s 3ms/step - loss: 3.4606 - mae: 0.6594 - mse: 0.6793 - val_loss: 3.1446 - val_mae: 0.4985 - val_mse: 0.4274

Edit: model creation code

def build_model(nhidden=5, nneurons=60, pdropout=.5,
                hidden_act='tanh', last_act='linear',
                loss='mae', regularizer=None, optimizer='adam',
                input_shape=None, weights=True):
    # Input: spectra (areas)
    x_input = Input(shape=input_shape)

    # Hidden layers
    hidden = x_input
    for i in range(nhidden):
        hidden = Dense(nneurons,
                       activation=hidden_act,
                       kernel_regularizer=regularizer,
                       name='dense{}'.format(i))(hidden)
        hidden = Dropout(pdropout)(hidden)

    # Last layer
    outputs = Dense(3, activation=last_act, name='denseout')(hidden)

    # Model
    model = Model(x_input, outputs)

    loss = my_mae
    model.compile(loss=my_mae,
                  optimizer=optimizer,
                  metrics=['mae', 'mse'])

    model.summary()

    return model
1 Answers

As you eventually found out, the culprit here is the use of a regularizer in your model:

hidden = Dense(nneurons,
               activation=hidden_act,
               kernel_regularizer=regularizer,   # <<<<<<<<<<
               name='dense{}'.format(i))(hidden)

Regularizers add an additional implicit loss, the regularization loss to your final loss value, which is (typically) independent of the specific data passed and based entirely on the regularization criteria chosen.

When you see the output in Keras, the loss value refers to the total loss, which is computed as the value returned by your loss function plus any regularization loss defined in the model.

Related