System information
- Have I written custom code (as opposed to using a stock example script provided in TensorFlow): Yes
- OS Platform and Distribution (e.g., Linux Ubuntu 16.04): Linux Ubuntu 16.04
- Mobile device (e.g. iPhone 8, Pixel 2, Samsung Galaxy) if the issue happens on mobile device:
- TensorFlow installed from (source or binary): pip
- TensorFlow version (use command below): 2.2
- Python version: 3.6.10
- Bazel version (if compiling from source):
- GCC/Compiler version (if compiling from source):
- CUDA/cuDNN version: CUDA 10.1, cuDNN 7.6.5
- GPU model and memory: GeForce GTX Titan X, 12 GB
Describe the current behavior
After migrating some Keras code from Theano backend (Keras 2.0.7, Theano 1.0.0) to pure TensorFlow, the model is unable to train with the same success as it did before. I did not change anything about the model architecture, hyper parameters, or parameters in the fit function. All that changed were import statements (ex: keras.layers -> tensorflow.keras.layers)
With TensorFlow:
2020-07-10 23:12:37 WARNING Using a generator with use_multiprocessing=True and multiple workers may duplicate your data. Please consider using the tf.data.Dataset.
Epoch 1/8
2020-07-10 23:12:38.795744: I tensorflow/stream_executor/platform/default/dso_loader.cc:44] Successfully opened dynamic library libcublas.so.10
2020-07-10 23:12:38 WARNING multiprocessing can interact badly with TensorFlow, causing nondeterministic deadlocks. For high performance data pipelines tf.data is recommended.
2020-07-10 23:12:39.211689: I tensorflow/stream_executor/platform/default/dso_loader.cc:44] Successfully opened dynamic library libcudnn.so.7
100/100 [==============================] - 10s 103ms/step - loss: 5.1734 - accuracy: 0.1002
Epoch 2/8
100/100 [==============================] - 10s 102ms/step - loss: 6.1694 - accuracy: 0.1079
Epoch 3/8
100/100 [==============================] - 10s 102ms/step - loss: 13.4283 - accuracy: 0.0830
Epoch 4/8
100/100 [==============================] - 10s 101ms/step - loss: 22.0164 - accuracy: 0.0838
Epoch 5/8
100/100 [==============================] - 10s 102ms/step - loss: 44.7066 - accuracy: 0.0575
Epoch 6/8
100/100 [==============================] - 10s 102ms/step - loss: 38.3269 - accuracy: 0.0690
Epoch 7/8
100/100 [==============================] - 10s 102ms/step - loss: 53.5934 - accuracy: 0.0641
Epoch 8/8
100/100 [==============================] - 10s 101ms/step - loss: 66.5487 - accuracy: 0.0673
Describe the expected behavior
Expected training behavior should be similar to that of Keras with Theano backend:
/usr/local/lib/python3.6/site-packages/keras/engine/training.py:1984: UserWarning: Using a generator with use_multiprocessing=True and multiple workers may duplicate your data. Please consider using the keras.utils.Sequence class.
UserWarning('Using a generator with use_multiprocessing=True')
Epoch 1/8
100/100 [==============================] - 14s - loss: 4.8412 - acc: 0.1197
Epoch 2/8
100/100 [==============================] - 13s - loss: 3.7391 - acc: 0.2454
Epoch 3/8
100/100 [==============================] - 13s - loss: 3.2523 - acc: 0.3528
Epoch 4/8
100/100 [==============================] - 13s - loss: 2.9821 - acc: 0.3923
Epoch 5/8
100/100 [==============================] - 13s - loss: 2.8871 - acc: 0.4059
Epoch 6/8
100/100 [==============================] - 13s - loss: 2.8087 - acc: 0.4177
Epoch 7/8
100/100 [==============================] - 13s - loss: 2.7203 - acc: 0.4290
Epoch 8/8
100/100 [==============================] - 13s - loss: 2.6507 - acc: 0.4351
Standalone code to reproduce the issue
Model architecture:
import tensorflow.keras.backend as K
from tensorflow.keras import Model, Sequential, Input
from tensorflow.keras.layers import Dense, Lambda, Embedding, LSTM
from tensorflow.keras.optimizers import Adam
from tensorflow.keras.losses import CategoricalCrossentropy
from tensorflow.keras.metrics import Accuracy
epoch = 1
batch_size = 2048
validation_batch_size = 1024
one_hot_dim = 80000
embedding_dim = 128
max_seq_len = 50
hidden_layer_dim = 512
max_q_size = 5000
nb_worker = 5
one_indexing = False
featurizer = "sequence"
featurizer_parameters = [one_hot_dim, max_seq_len, one_indexing]
def create_model_architecture(num_classes):
# build model architecture
model = Sequential(
[
Input(shape=(max_seq_len,), name="query_input"),
Embedding(output_dim=embedding_dim, input_dim=one_hot_dim, input_length=max_seq_len),
LSTM(embedding_dim, input_shape=(max_seq_len, embedding_dim)),
Lambda(lambda x: K.cast(x, "float32")),
Dense(units=hidden_layer_dim, kernel_initializer="he_normal", activation="relu"),
Dense(units=num_classes, kernel_initializer="normal", activation="softmax", name="prediction")
]
)
model.compile(optimizer='adam', loss='categorical_crossentropy', metrics=['accuracy']
model.summary()
return model
In model.compile, I have tried also using the classes for optimizer, loss, and metrics (such as Adam, CategoricalCrossentropy, Accuracy). It causes the accuracy to be much higher but the loss still increases exponentially.
Model fit call:
model.fit(data_generator_train.generate(),
steps_per_epoch=args.step_log,
epochs=8,
validation_data=data_generator_validate.generate() if data_generator_validate else None,
validation_steps=data_generator_validate.num_batches if data_generator_validate else None,
max_queue_size=model_module.max_q_size,
shuffle=True,
use_multiprocessing=True,
workers=model_module.nb_worker,
callbacks=[check_pointer])
Data generator function:
@threadsafe_generator
def generate(self):
while True:
for filename in glob.glob(self.dataset_files):
with open(filename) as fr:
while True:
lines = list(islice(fr, self.batch_size))
if not lines:
break
x, y, w = self.get_train_vectors(lines)
yield x, y, w
Other info / logs
Other things I have tried:
Using tensorflow.keras.utils.Sequence instead of a generator, and including shuffle=True Not using multiprocessing Not using Sequential Using classes instead of names in model.compile() Changing parameters like batch size, steps_per_epoch, etc.
Other notes:
I am using a custom Docker image to run the model. I have made volume mountings from the local machine, including CUDA, cuDNN, etc. under /usr/local/cuda. TensorFlow 2.2 also has a library (libcublas.so) which is not located there, so I have an additional mapping under /usr/lib/x86_64-linux-gnu.