I train a model with tensorflow.keras on two GPUs, Tesla M60.
The model gets trained successfully when I limit the amount of training-data below a certain value (around 700). When i go above that value I get an out of memory error for the GPU (overall training-data size is 2196).
I use on Windows Server 2016:
(I tried to fulfill this https://www.tensorflow.org/install/source_windows?hl=en)
- python 3.9.7
- tensorflow 2.6.0
- tensorflow-gpu 2.6.0
- cudatoolkit 11.2.2
- cudnn 8.1.0.77
If the exact build is neccessary, please tell me.
I use the same setting for the model- and training configuration, but feed different size of training-data:
learning_rate = 0.001
epochs = 200
batch_size = 80
buffer = 1000;
x_train = io.loadmat(...)
x_train = x_train[0:700,:,:] # optional
y_train = io.loadmat(...)
y_train = x_train[0:700,:,:] # optional
x_test = io.loadmat(...)
y_test = io.loadmat(...)
strategy = tensorflow.distribute.MirroredStrategy(cross_device_ops=tensorflow.distribute.HierarchicalCopyAllReduce())
with strategy.scope():
model= tensorflow.keras.Model(...)
loss = tensorflow.keras.losses.MeanSquaredError(name='MeanSquaredError')
model.compile(optimizer=tensorflow.keras.optimizers.Adam(learning_rate=learning_rate), loss=loss])
train_data = tensorflow.data.Dataset.from_tensor_slices((x_train, y_train))
val_data = tensorflow.data.Dataset.from_tensor_slices((x_test, y_test))
train_data = train_data.shuffle(buffer).batch(batch_size)
val_data = val_data.shuffle(buffer).batch(batch_size)
options = tensorflow.data.Options()
options.experimental_distribute.auto_shard_policy = tensorflow.data.experimental.AutoShardPolicy.DATA
train_data = train_data.with_options(options)
val_data = val_data.with_options(options)
model.fit(train_data,
validation_data=val_data,
epochs=epochs,
callbacks=[livelossplot.PlotLossesKeras(outputs=[livelossplot.outputs.MatplotlibPlot(figpath=save_string_plot)]),tensorflow.keras.callbacks.TensorBoard(log_dir="model\\" + string_name_model, histogram_freq = True, write_grads=True)],
verbose=0)
When I do not use tensorflow.data it works for a size until 1100 and then crashes (with the warning: AUTO sharding policy will apply DATA sharding policy as it failed to apply FILE sharding policy because of the following reason: Did not find a shardable source, walked to a node which is not a dataset):
model.fit(x_train, y_train,
validation_data=(x_test, y_test),
epochs=epochs,
batch_size= batch_size,
shuffle = 1,
callbacks=[livelossplot.PlotLossesKeras(outputs=[livelossplot.outputs.MatplotlibPlot(figpath=save_string_plot)]),tensorflow.keras.callbacks.TensorBoard(log_dir="model\\" + string_name_model, histogram_freq = True, write_grads=True)],
verbose=0)
This confuses me, as I think the GPU memory usage is mainly dependent on the batch size, not the amount of training-data and for smaller training-sizes it works with the same batch size.
Can anyone tell me, what is wrong in my thinking?
EDIT:
I tested the exact same code for another configuration of python installation:
- python 3.6.12
- tensorflow 2.1.0
- tensorflow-gpu 2.1.0
- cudatoolkit 10.1.243
- cudnn 7.6.5
And now the amount of training-data (now over 3000) does not affect the GPU memory, so no out of memory any more (each case with tensorflow.data and without).
I wanted to use the latest version because of some additonal features. So can anyone tell me still, what might went wrong?