GAN training and unexpected NaN results

Viewed 441

Hi everyone and happy 2021.

I am trying to train a GAN network, being the input data a dataset with 3 columns (Each with values between -1 and 1), into a NumPy array shape (20000,3), looking like these 3 examples:

 [0.7758508  0.99201597 0.90802954]
 [0.99263527 0.99600798 0.90231117]]

However, very early in the training, all losses start to go to NaN, like below. The generator output is also going to NaN

GEN Loss:  0.019527255
DISC Loss:  [-0.055139657, -0.07087045, -0.017024973, 0.032755766]
GEN Loss:  0.017767208
DISC Loss:  [-0.060895607, -0.070055336, -0.01139014, 0.02054987]
GEN Loss:  0.015142502
DISC Loss:  [-0.019746825, -0.06709217, -0.017383143, 0.06472849]
GEN Loss:  0.016435547
DISC Loss:  [-0.037867796, -0.07073814, -0.018058099, 0.05092844]
GEN Loss:  0.019013882
DISC Loss:  [nan, nan, -0.013940675, 0.05671501]
GEN Loss:  nan
DISC Loss:  [nan, nan, nan, nan]

I decided to try with the WGAN Keras implementation at https://github.com/keras-team/keras-contrib/blob/master/examples/improved_wgan.py, and still I have the same behavior.

What could be the root cause?

See below model and loss functions definitions (losses functions are as in the WGAN link above):

def make_discriminator():
    
    model = Sequential()
    model.add(Dense(25, input_dim=3))
    model.add(LeakyReLU())
    model.add(Dense(40))
    model.add(LeakyReLU())
    model.add(Dense(24, kernel_initializer='he_normal'))
    model.add(LeakyReLU())
    model.add(Dense(1, kernel_initializer='he_normal'))
    return model


def make_generator():
    """Creates a generator model that takes a 20-dimensional noise vector as a "seed",
    and outputs images of size 28x28x1."""
    model = Sequential()
    model.add(Dense(50, input_dim=20))
    model.add(BatchNormalization())
    model.add(LeakyReLU())
    model.add(Dense(3,  activation='tanh'))
   return model
discriminator_model = Model(inputs=[real_samples, generator_input_for_discriminator],
                            outputs=[discriminator_output_from_real_samples,discriminator_output_from_generator, averaged_samples_out])

discriminator_model.compile(optimizer=Adam(0.0001, beta_1=0.5, beta_2=0.9),
                            loss=[wasserstein_loss, wasserstein_loss,partial_gp_loss])

All clues are really appreciated.

0 Answers
Related