Switching from Theano 1.0.0 to TensorFlow 2.2 (both using Keras) causes exponentially growing loss and stagnant accuracy

Viewed 55

System information

  • Have I written custom code (as opposed to using a stock example script provided in TensorFlow): Yes
  • OS Platform and Distribution (e.g., Linux Ubuntu 16.04): Linux Ubuntu 16.04
  • Mobile device (e.g. iPhone 8, Pixel 2, Samsung Galaxy) if the issue happens on mobile device:
  • TensorFlow installed from (source or binary): pip
  • TensorFlow version (use command below): 2.2
  • Python version: 3.6.10
  • Bazel version (if compiling from source):
  • GCC/Compiler version (if compiling from source):
  • CUDA/cuDNN version: CUDA 10.1, cuDNN 7.6.5
  • GPU model and memory: GeForce GTX Titan X, 12 GB

Describe the current behavior

After migrating some Keras code from Theano backend (Keras 2.0.7, Theano 1.0.0) to pure TensorFlow, the model is unable to train with the same success as it did before. I did not change anything about the model architecture, hyper parameters, or parameters in the fit function. All that changed were import statements (ex: keras.layers -> tensorflow.keras.layers)

With TensorFlow:

2020-07-10 23:12:37 WARNING  Using a generator with use_multiprocessing=True and multiple workers may duplicate your data. Please consider using the tf.data.Dataset.
Epoch 1/8
2020-07-10 23:12:38.795744: I tensorflow/stream_executor/platform/default/dso_loader.cc:44] Successfully opened dynamic library libcublas.so.10
2020-07-10 23:12:38 WARNING  multiprocessing can interact badly with TensorFlow, causing nondeterministic deadlocks. For high performance data pipelines tf.data is recommended.
2020-07-10 23:12:39.211689: I tensorflow/stream_executor/platform/default/dso_loader.cc:44] Successfully opened dynamic library libcudnn.so.7
100/100 [==============================] - 10s 103ms/step - loss: 5.1734 - accuracy: 0.1002
Epoch 2/8
100/100 [==============================] - 10s 102ms/step - loss: 6.1694 - accuracy: 0.1079
Epoch 3/8
100/100 [==============================] - 10s 102ms/step - loss: 13.4283 - accuracy: 0.0830
Epoch 4/8
100/100 [==============================] - 10s 101ms/step - loss: 22.0164 - accuracy: 0.0838
Epoch 5/8
100/100 [==============================] - 10s 102ms/step - loss: 44.7066 - accuracy: 0.0575
Epoch 6/8
100/100 [==============================] - 10s 102ms/step - loss: 38.3269 - accuracy: 0.0690
Epoch 7/8
100/100 [==============================] - 10s 102ms/step - loss: 53.5934 - accuracy: 0.0641
Epoch 8/8
100/100 [==============================] - 10s 101ms/step - loss: 66.5487 - accuracy: 0.0673

Describe the expected behavior

Expected training behavior should be similar to that of Keras with Theano backend:

/usr/local/lib/python3.6/site-packages/keras/engine/training.py:1984: UserWarning: Using a generator with use_multiprocessing=True and multiple workers may duplicate your data. Please consider using the keras.utils.Sequence class.
  UserWarning('Using a generator with use_multiprocessing=True')
Epoch 1/8
100/100 [==============================] - 14s - loss: 4.8412 - acc: 0.1197          
Epoch 2/8
100/100 [==============================] - 13s - loss: 3.7391 - acc: 0.2454     
Epoch 3/8
100/100 [==============================] - 13s - loss: 3.2523 - acc: 0.3528     
Epoch 4/8
100/100 [==============================] - 13s - loss: 2.9821 - acc: 0.3923     
Epoch 5/8
100/100 [==============================] - 13s - loss: 2.8871 - acc: 0.4059     
Epoch 6/8
100/100 [==============================] - 13s - loss: 2.8087 - acc: 0.4177     
Epoch 7/8
100/100 [==============================] - 13s - loss: 2.7203 - acc: 0.4290     
Epoch 8/8
100/100 [==============================] - 13s - loss: 2.6507 - acc: 0.4351

Standalone code to reproduce the issue

Model architecture:

import tensorflow.keras.backend as K
from tensorflow.keras import Model, Sequential, Input
from tensorflow.keras.layers import Dense, Lambda, Embedding, LSTM
from tensorflow.keras.optimizers import Adam
from tensorflow.keras.losses import CategoricalCrossentropy
from tensorflow.keras.metrics import Accuracy

epoch = 1
batch_size = 2048
validation_batch_size = 1024
one_hot_dim = 80000
embedding_dim = 128
max_seq_len = 50
hidden_layer_dim = 512
max_q_size = 5000

nb_worker = 5
one_indexing = False
featurizer = "sequence"
featurizer_parameters = [one_hot_dim, max_seq_len, one_indexing]
def create_model_architecture(num_classes):
    # build model architecture
    model = Sequential(
         [
             Input(shape=(max_seq_len,), name="query_input"),
             Embedding(output_dim=embedding_dim, input_dim=one_hot_dim, input_length=max_seq_len),
             LSTM(embedding_dim, input_shape=(max_seq_len, embedding_dim)),
             Lambda(lambda x: K.cast(x, "float32")),
             Dense(units=hidden_layer_dim, kernel_initializer="he_normal", activation="relu"),
             Dense(units=num_classes, kernel_initializer="normal", activation="softmax", name="prediction")
         ]
    )
    model.compile(optimizer='adam', loss='categorical_crossentropy', metrics=['accuracy']
    model.summary()
    return model

In model.compile, I have tried also using the classes for optimizer, loss, and metrics (such as Adam, CategoricalCrossentropy, Accuracy). It causes the accuracy to be much higher but the loss still increases exponentially.

Model fit call:

model.fit(data_generator_train.generate(),
            steps_per_epoch=args.step_log,
            epochs=8,
            validation_data=data_generator_validate.generate() if data_generator_validate else None,
            validation_steps=data_generator_validate.num_batches if data_generator_validate else None,
            max_queue_size=model_module.max_q_size,
            shuffle=True,
            use_multiprocessing=True,
            workers=model_module.nb_worker,
            callbacks=[check_pointer])

Data generator function:

@threadsafe_generator
def generate(self):
   while True:
      for filename in glob.glob(self.dataset_files):
         with open(filename) as fr:
            while True:
               lines = list(islice(fr, self.batch_size))
               if not lines:
                  break
               x, y, w = self.get_train_vectors(lines)
               yield x, y, w

Other info / logs

Other things I have tried:

Using tensorflow.keras.utils.Sequence instead of a generator, and including shuffle=True Not using multiprocessing Not using Sequential Using classes instead of names in model.compile() Changing parameters like batch size, steps_per_epoch, etc.

Other notes:

I am using a custom Docker image to run the model. I have made volume mountings from the local machine, including CUDA, cuDNN, etc. under /usr/local/cuda. TensorFlow 2.2 also has a library (libcublas.so) which is not located there, so I have an additional mapping under /usr/lib/x86_64-linux-gnu.

0 Answers
Related