How to print the maximum memory used during Keras's model.fit()

Viewed 364

I wrote a neural network model using Keras and Tensorflow and was able to train and run it. At this point, I want to know how much memory was required for training the model. How can I print this information during the training phase? I tried the Keras model profiler below but it didn't explain the peak memory required for the training phase. For example, training my model shows out of memory on 6GB GPU card but the profile says that the memory requirement is less than 1GB. So, how can I measure the peak run-time memory requirement when I use model.fit() in Keras?

https://github.com/Mr-TalhaIlyas/Tensorflow-Keras-Model-Profiler

1 Answers

I would suggest using a Keras Callback and printing the GPU usage after every epoch for example. You can get the GPU information with tf.config.experimental.get_memory_info('GPU:0'). Here is a working example:

import tensorflow as tf

class MemoryPrintingCallback(tf.keras.callbacks.Callback):
    def on_epoch_end(self, epoch, logs=None):
      gpu_dict = tf.config.experimental.get_memory_info('GPU:0')
      tf.print('\n GPU memory details [current: {} gb, peak: {} gb]'.format(
          float(gpu_dict['current']) / (1024 ** 3), 
          float(gpu_dict['peak']) / (1024 ** 3)))
        
inputs = tf.keras.layers.Input((1000,))
x = tf.keras.layers.Dense(1000, 'relu')(inputs)
x = tf.keras.layers.Dense(1000, 'relu')(x)
x = tf.keras.layers.Dense(1000, 'relu')(x)
x = tf.keras.layers.Dense(1000, 'relu')(x)
outputs = tf.keras.layers.Dense(1, 'sigmoid')(x)
model = tf.keras.Model(inputs, outputs)
model.compile(optimizer='adam', loss=tf.keras.losses.BinaryCrossentropy())

x = tf.random.normal((500, 1000))
y = tf.random.uniform((500, 1), maxval=2, dtype=tf.int32)
model.fit(x, y, batch_size=50, epochs = 20, callbacks= [MemoryPrintingCallback()])
 GPU memory details [current: 0.321030855178833 gb, peak: 0.32660841941833496 gb]
Epoch 1/20
10/10 [==============================] - 1s 8ms/step - loss: 0.9309

 GPU memory details [current: 0.3508758544921875 gb, peak: 0.3557243347167969 gb]
Epoch 2/20
10/10 [==============================] - 0s 7ms/step - loss: 0.5702

 GPU memory details [current: 0.3508758544921875 gb, peak: 0.3557243347167969 gb]
Epoch 3/20
10/10 [==============================] - 0s 8ms/step - loss: 0.1311

 GPU memory details [current: 0.3508758544921875 gb, peak: 0.3557243347167969 gb]
Epoch 4/20
10/10 [==============================] - 0s 7ms/step - loss: 0.0865

 GPU memory details [current: 0.3508758544921875 gb, peak: 0.3661658763885498 gb]
Epoch 5/20
10/10 [==============================] - 0s 7ms/step - loss: 0.0379
...

You can find out the name of your device with:

print(tf.config.list_physical_devices('GPU'))
#[PhysicalDevice(name='/physical_device:GPU:0', device_type='GPU')]

Note the following though:

For GPUs, TensorFlow will allocate all the memory by default, unless changed with tf.config.experimental.set_memory_growth. The dict specifies only the current and peak memory that TensorFlow is actually using, not the memory that TensorFlow has allocated on the GPU. source

Related