GPU Out of Memory for different training-data sizes but same configurations (batch_size etc.), tensorflow.keras

Viewed 162

I train a model with tensorflow.keras on two GPUs, Tesla M60.

The model gets trained successfully when I limit the amount of training-data below a certain value (around 700). When i go above that value I get an out of memory error for the GPU (overall training-data size is 2196).

I use on Windows Server 2016:

(I tried to fulfill this https://www.tensorflow.org/install/source_windows?hl=en)

  • python 3.9.7
  • tensorflow 2.6.0
  • tensorflow-gpu 2.6.0
  • cudatoolkit 11.2.2
  • cudnn 8.1.0.77

If the exact build is neccessary, please tell me.

I use the same setting for the model- and training configuration, but feed different size of training-data:

learning_rate = 0.001
epochs = 200
batch_size = 80
buffer = 1000;

x_train = io.loadmat(...)
x_train = x_train[0:700,:,:] # optional
y_train = io.loadmat(...)
y_train = x_train[0:700,:,:] # optional
x_test = io.loadmat(...)
y_test = io.loadmat(...)

strategy = tensorflow.distribute.MirroredStrategy(cross_device_ops=tensorflow.distribute.HierarchicalCopyAllReduce())

with strategy.scope():
    
     model= tensorflow.keras.Model(...)
    
     loss = tensorflow.keras.losses.MeanSquaredError(name='MeanSquaredError')

     model.compile(optimizer=tensorflow.keras.optimizers.Adam(learning_rate=learning_rate), loss=loss])

train_data = tensorflow.data.Dataset.from_tensor_slices((x_train, y_train))
val_data = tensorflow.data.Dataset.from_tensor_slices((x_test, y_test))

train_data = train_data.shuffle(buffer).batch(batch_size)
val_data = val_data.shuffle(buffer).batch(batch_size)

options = tensorflow.data.Options()
options.experimental_distribute.auto_shard_policy = tensorflow.data.experimental.AutoShardPolicy.DATA

train_data = train_data.with_options(options)
val_data = val_data.with_options(options)

model.fit(train_data,
          validation_data=val_data,
          epochs=epochs,
          callbacks=[livelossplot.PlotLossesKeras(outputs=[livelossplot.outputs.MatplotlibPlot(figpath=save_string_plot)]),tensorflow.keras.callbacks.TensorBoard(log_dir="model\\" + string_name_model, histogram_freq = True, write_grads=True)],
          verbose=0)

When I do not use tensorflow.data it works for a size until 1100 and then crashes (with the warning: AUTO sharding policy will apply DATA sharding policy as it failed to apply FILE sharding policy because of the following reason: Did not find a shardable source, walked to a node which is not a dataset):

model.fit(x_train, y_train,
          validation_data=(x_test, y_test),
          epochs=epochs,
          batch_size= batch_size,
          shuffle = 1,
          callbacks=[livelossplot.PlotLossesKeras(outputs=[livelossplot.outputs.MatplotlibPlot(figpath=save_string_plot)]),tensorflow.keras.callbacks.TensorBoard(log_dir="model\\" + string_name_model, histogram_freq = True, write_grads=True)],
          verbose=0)

This confuses me, as I think the GPU memory usage is mainly dependent on the batch size, not the amount of training-data and for smaller training-sizes it works with the same batch size.

Can anyone tell me, what is wrong in my thinking?

EDIT:

I tested the exact same code for another configuration of python installation:

  • python 3.6.12
  • tensorflow 2.1.0
  • tensorflow-gpu 2.1.0
  • cudatoolkit 10.1.243
  • cudnn 7.6.5

And now the amount of training-data (now over 3000) does not affect the GPU memory, so no out of memory any more (each case with tensorflow.data and without).

I wanted to use the latest version because of some additonal features. So can anyone tell me still, what might went wrong?

0 Answers
Related