How much memory does Keras allocate while adding a layer to sequential model?

Viewed 94

I am using Colab GPU to code a research paper's architecture using Keras. I get a resourceexhaustederror while adding the first dense layer to the sequential model. The following are the outputs of nvidia-smi before and after the error -

After loading dataset and before creating model -

+-----------------------------------------------------------------------------+
| NVIDIA-SMI 460.67       Driver Version: 460.32.03    CUDA Version: 11.2     |
|-------------------------------+----------------------+----------------------+
| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |
| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |
|                               |                      |               MIG M. |
|===============================+======================+======================|
|   0  Tesla K80           Off  | 00000000:00:04.0 Off |                    0 |
| N/A   43C    P0    59W / 149W |   4238MiB / 11441MiB |      0%      Default |
|                               |                      |                  N/A |
+-------------------------------+----------------------+----------------------+
                                                                               
+-----------------------------------------------------------------------------+
| Processes:                                                                  |
|  GPU   GI   CI        PID   Type   Process name                  GPU Memory |
|        ID   ID                                                   Usage      |
|=============================================================================|
+-----------------------------------------------------------------------------+

Error while adding first layer in model -

fc2_shape = 128*128

model = models.Sequential()
model.add(layers.Flatten(input_shape=(128, 128, 2)))
# with tf.device('/gpu:0'):
model.add(layers.Dense(fc2_shape*2, activation='tanh'))
# with tf.device('/gpu:1'):
model.add(layers.Dense(fc2_shape, activation='tanh'))
model.add(layers.Dense(fc2_shape, activation='tanh'))
model.add(layers.Reshape((128, 128, 1)))
model.add(layers.Conv2D(64, (5, 5), activation='relu'))
model.add(layers.Conv2D(64, (5, 5), activation='relu', kernel_regularizer=regularizers.l1(0.0001)))
model.add(layers.Conv2DTranspose(1, kernel_size=9, strides=1))

model.summary()

---------------------------------------------------------------------------
ResourceExhaustedError                    Traceback (most recent call last)
<ipython-input-12-da3c18fe6dbe> in <module>()
      4 model.add(layers.Flatten(input_shape=(128, 128, 2)))
      5 # with tf.device('/gpu:0'):
----> 6 model.add(layers.Dense(fc2_shape*2, activation='tanh'))
      7 # with tf.device('/gpu:1'):
      8 model.add(layers.Dense(fc2_shape, activation='tanh'))

25 frames
/usr/local/lib/python3.7/dist-packages/six.py in raise_from(value, from_value)

ResourceExhaustedError: OOM when allocating tensor with shape[32768,32768] and type float on /job:localhost/replica:0/task:0/device:GPU:0 by allocator GPU_0_bfc [Op:Add]

Output of nvidia-smi after error -

+-----------------------------------------------------------------------------+
| NVIDIA-SMI 460.32.03    Driver Version: 460.32.03    CUDA Version: 11.2     |
|-------------------------------+----------------------+----------------------+
| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |
| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |
|                               |                      |               MIG M. |
|===============================+======================+======================|
|   0  Tesla T4            Off  | 00000000:00:04.0 Off |                    0 |
| N/A   61C    P0    30W /  70W |  14276MiB / 15109MiB |      0%      Default |
|                               |                      |                  N/A |
+-------------------------------+----------------------+----------------------+
                                                                               
+-----------------------------------------------------------------------------+
| Processes:                                                                  |
|  GPU   GI   CI        PID   Type   Process name                  GPU Memory |
|        ID   ID                                                   Usage      |
|=============================================================================|
+-----------------------------------------------------------------------------+

A float tensor of shape [32678, 32768] will need space of 32768 x 32768 x 4 bytes or 4.29 GB. Why do I get an oom error when I have sufficient memory available for 4.29 GB? From ~4.44 GB (4238 MiB) allocated before the error to ~14.96 GB allocated after the error, Keras seems to be allocating around 10 GB memory for this layer. I wish to know what all for does Keras precisely allocate memory (maybe weights, weight gradients, input gradients etc.) while adding a layer, so that I can estimate my GPU memory requirement for a given task. Thanks.

0 Answers
Related