Keras multilayer custom loss function equals zero

Viewed 81

I have an autoencoder and I am trying to implement the third solution from this article. So I have one hot encoded columns and I want to use multitask learning to optimize all dummy variables (say country_US, country_NL and country_UK) belonging to one original column (country) separately. In this way the separate layers represent multiclass classification problems. I created a custom loss function for this which seems to work. This is my code, including example dataset:

from keras.models import Model
from keras.layers import Dense, Input
from keras.optimizers import Adam
from keras.callbacks import EarlyStopping
import numpy as np
import pandas as pd
import random
import keras
from collections import Counter
from keras.layers.merge import concatenate
from sklearn.model_selection import train_test_split

data = pd.DataFrame({'country':random.choices(['UK', 'US', 'NL', 'DE', 'BE'],k=1000),
                     'period': random.choices([1,2,3,4,5,6,7,8,9,10,11,12],k=1000),
                     'fruit': random.choices(['apple', 'mango', 'banana','cherry','grape','strawberry','melon'], k=1000),
                     'shoes': random.choices(['yes', 'no'],k=1000),
                     'something': random.choices(['a', 'b','c', 'd'], k=1000),
                     'sports': random.choices(['soccer', 'hockey', 'handball', 'tennis', 'dodgeball'], k=1000)})
data = pd.get_dummies(data.astype(str))
train, vali = train_test_split(data, test_size=0.1, random_state=42)

def autoencoder_mult(train, vali, columns, dimensions, activation, optimizer, batch_size, epochs, early_stop, alpha=0.01):   
    input_shape = train.shape[1]
    column_types = Counter([col.split('_')[0] for col in columns])
    input_layers = []
    for col in column_types.keys():
        input_layers.append(Input(shape=(column_types[col],), name=col+'_in'))
    inputs = concatenate(input_layers)    
    hidden = Dense(dimensions[0], activation=activation)(inputs)
    for i in range(1,len(dimensions)):
        hidden = Dense(dimensions[i], activation=activation)(hidden)
    hidden = Dense(input_shape, activation=activation)(hidden)
    output_layers=[]
    losses = {}
    y = []
    validation_data=[]
    for col in column_types.keys():
        output_layers.append(Dense(column_types[col], activation='softmax', name=col)(hidden))
        subset = [columns.get_loc(c) for c in columns if col in c]
        losses[col] = custom_loss(subset)
        y.append(train[:,subset])
        validation_data.append(vali[:,subset])
    model = Model(inputs = input_layers, outputs = output_layers)
    model.compile(optimizer = optimizer, loss = losses,metrics = ['mse','accuracy']) 
    history = model.fit(x = y, y = y, batch_size = batch_size, shuffle = True,
              epochs = epochs, verbose = 2,  callbacks = [early_stop], validation_data=(validation_data,validation_data)) 
    return model, history    

def custom_loss(subset):
    def loss_fn(true, pred):
        return keras.losses.categorical_crossentropy(true[:,subset[0]:(subset[-1]+1)],pred[:,subset[0]:(subset[-1]+1)])
    return loss_fn

def np_int(x):
    return np.asarray(x).astype(np.int64)


model,history = autoencoder_mult(np_int(train), np_int(vali), train.columns, [20,10,20], 'relu', Adam(lr = 0.0001), 128, 200, EarlyStopping(monitor='val_loss', min_delta=0.0001, patience=5))

However, the autoencoder is not working well. This is an example of my output during one epoch:

loss: 1.1162 - country_loss: 0.2209 - period_loss: 0.8954 - fruit_loss: 0.0000e+00 - shoes_loss: 0.0000e+00 - something_loss: 0.0000e+00 - sports_loss: 0.0000e+00 - country_mse: 0.0155 - country_accuracy: 0.9889 - period_mse: 0.0838 - period_accuracy: 0.1044 - fruit_mse: 0.1494 - fruit_accuracy: 0.1778 - shoes_mse: 0.2863 - shoes_accuracy: 0.5022 - something_mse: 0.2086 - something_accuracy: 0.2567 - sports_mse: 0.1746 - sports_accuracy: 0.1911 - 
val_loss: 1.1920 - val_country_loss: 0.2417 - val_period_loss: 0.9503 - val_fruit_loss: 0.0000e+00 - val_shoes_loss: 0.0000e+00 - val_something_loss: 0.0000e+00 - val_sports_loss: 0.0000e+00 - val_country_mse: 0.0187 - val_country_accuracy: 0.9800 - val_period_mse: 0.0825 - val_period_accuracy: 0.1400 - val_fruit_mse: 0.1582 - val_fruit_accuracy: 0.1000 - val_shoes_mse: 0.3186 - val_shoes_accuracy: 0.4600 - val_something_mse: 0.2156 - val_something_accuracy: 0.2200 - val_sports_mse: 0.1751 - val_sports_accuracy: 0.1700

Only for the first two columns (in my actual dataset actually only the first column) it calculates a loss, for all other columns the loss equals zero. It also stops after a few iterations due to the early stopping criterion, it doesn't learn much. Anyone have an idea why this happens? Thanks!

0 Answers
Related