Custom loss function to prevent asymmetrical confusion matrices

Viewed 98

I am trying to write a custom loss function in Tensorflow/Keras that will add a penalty to the loss function if the number ratio between positive predicted classes is not identical (asymmetrical confusion matrix).

The confusion matrix for one of my models is for example:

tf.Tensor(
[[8004 2908]
 [7860 3063]], shape=(2, 2), dtype=int32)

The solving process was stopped after a couple of epochs before the train_loss got worse. While the confusion matrix shows a majority of positive predicted classes on the main diagonal already the overall matrix is asymmetrical. My "hope" is to improve the training process via a custom loss function. Basically i want to manipulate the form of the confusion matrix during the training. I came up with the following function (here for 4 classes that are treated as 2):

def sym_loss(y_true, y_pred):
    index_list = tf.cast(tf.math.argmax(y_pred, axis=1), tf.float32)
    conditions = tf.greater(index_list, 1)
    mean_index = tf.reduce_mean(index_list)
    if_value = tf.math.subtract(mean_index, 1.5) #1.5 hard coded as mean index for 4 classes
    corr_value = tf.abs(if_value)

    if if_value < 0:
        mask_tensor = tf.where(conditions, corr_value, 0.0)
    else:
        mask_tensor = tf.where(conditions, 0.0, corr_value)

    squared_difference = tf.square(tf.cast(y_true, tf.float32) - tf.cast(y_pred, tf.float32))
    solu = tf.reduce_mean(squared_difference, axis=-1)

    return tf.math.add(solu, mask_tensor)

The function itself runs without errors but only improve the shape of the confusion matrix sometimes.

The idea behind the loss function is to evaluate the ratio of predictions that are in one or another class. Ideally the ratio should be 1. If a miss match is detected the overall distance " (here corr_value) is calculated. Than i calculate one arbitrary standard loss function (here mean square difference) and add the overall distance only to the predicted class that is dominant. In the shown example two of the four classes are considered as equal. A optimization for a true binary loss or general loss function would be possible. I hope it is not to confusing what i want to achieve.

The problem is that the returned loss function is not a scalar but a tensor with the size of samples in one batch (as far as i know). My guess is that this kind of penalization will be ignored by the optimizer and just acts as additional noise. When i train my models i can see that the loss function is more "noisy" than usual.

Another idea would be to directly evaluate the current confusion matrix within the loss function and manipulate the returned tensor. But than i would still have the problem how to perform the calculation correctly.

Does anyone have an idea how to address this problem properly?

Thanks in advance!

EDIT:

I searched a little bit and found some publications regarding this problem. It seems my problem is generally described as an "biased classifier".

A similar problem can be found here: https://towardsdatascience.com/understanding-and-reducing-bias-in-machine-learning-6565e23900ac

Still i haven't found an option to implement this into keras/tensorflow. Maybe there is already something implemented in tensorflow to address this biased classifier problem but i am not able to find it...

0 Answers
Related