Reducing Losses of Autoencoder

Viewed 1665

i am currently trying to train an autoencoder which allows the representation of an array with the length of 128 integer variables to a compression of 64. The array contains 128 integer values ranging from 0 to 255.

I train the model with over 2 million datapoints each epoch. Each array has a form like this: [ 1, 9, 0, 4, 255, 7, 6, ..., 200]

input_img = Input(shape=(128,))
encoded = Dense(128, activation=activation)(input_img)
encoded = Dense(128, activation=activation)(encoded)

encoded = Dense(64, activation=activation)(encoded)

decoded = Dense(128, activation=activation)(encoded)
decoded = Dense(128, activation='linear')(decoded)

autoencoder = Model(input_img, decoded)
autoencoder.compile(optimizer='adam', loss='mse')

history = autoencoder.fit(np.array(training), np.array(training),
                    epochs=50,
                    batch_size=256,
                    shuffle=True,
                    validation_data=(np.array(test), np.array(test)),
                    callbacks=[checkpoint, early_stopping])

I will also upload a graphic showing the training and validation process: Loss graph of Training

How is it possible for me to lower the loss further. What I have tried so far (neither option has led to success):

  1. Longer training phase
  2. More Layer
1 Answers

There is of course not a magic thing that you can do to instantly reduce the loss as it is very problem specific, but here is a couple tricks that I could suggest:

  • Reduce mini-batch size. Having a smaller batch size will make the gradient more noisy when it's back-propagating. This might seem counter-intuitive first, but this noise in the gradient descent could help the descent overcome possible local minimas. Think of it this way; when the descent is noisy, it will take longer but the plateau will be lower, when the descent is smooth, it will take less but will settle in an earlier plateau. (Very generalized!)
  • Try to make the layers have units with expanding/shrinking order. So instead of using 128 unit layers back to back, make it 128 to 256. This way, you wouldn't be forcing the model to represent 128 numbers with another pack of 128 numbers. You could have all the layers with 128 units, that would in theory result with a lossless autoencoding, where the input and the output is literally the same. But that does not happen in practice, becacuse of the nature of the gradient descent. It's like you randomly start somewhere in the jungle and try to make your way through it following a lead (negative gradient) but just because you have the lead does not mean you can reach where you are headed. So to get meaningful information out of your distribution, you should force your model to represent information with lesser units. This will make the work of gradient descent easier, because you are setting a prior condition; if it can't encode the information well enough, it will have a high loss. So you kinda make it understand what you desire from the model.
  • The absolute value of the error function. You are trying to lower your loss, but to what end? Do you need it to go near 0, or do you just need it to be lower as possible? Because as your latent dimension shrinks, the loss will increase but the autoencoder will be able to capture the latent representative information of the data better. Because you are forcing the encoder to represent an information of higher dimension with an information with lower dimension. So the lower the latent dimension is, the more the autoencoder will try working toward extracting the most meaningful information out of the input, because it has limited space. So, even the loss is greater the distribution is captured more competently. So it is up to your problem, if you want something like noise reduction on images, go with higher encoding dimensions, but if you want to do something like anomaly detection, it's better to try lower dimensions without completely gutting the representative capacity of the model.
  • This is a bit more tinfoil advice of mine but you also try to shift your numbers down so that the range is -128 to 128. I have -not so accurately- observed that some activations (especially ReLU) work slightly better with these kind of inputs.

I hope some of these works for you. Good luck.

Related