Memory leaking after `tf.keras.Model.fit` is called and training doesn't start

Viewed 262

I'm using my yolo implementation which used to work fine on tensorflow versions prior to 2.5. I tried recently training yolo3 on a small dataset (which uses tf.keras.Model.fit). Here's a colab notebook which you can use to reproduce the issue. Shortly after model.fit is called, the messages below keep repeating:

/usr/local/lib/python3.7/dist-packages/keras/engine/functional.py:1410: CustomMaskWarning: Custom mask layers require a config and must override get_config. When loading, the custom mask layer must be passed to the custom_objects argument.
  layer_config = serialize_layer_fn(layer)

and

INFO:tensorflow:Assets written to: ram://eefa3127-ad7d-4445-a186-75fd8f0b81e1/assets

Then memory usage keeps growing for no apparent reason and eventually a memory crash occurs. (which doesn't happen in earlier tensorflow versions <= 2.5). You can verify so using this other notebook which uses tensorflow 2.5 instead, things should go perfectly fine and training goes as expected. I also tried installing tensorflow 2.8 instead of colab's default version (2.7) and the issue persists.

Here's the output containing problems (tensorflow > 2.5):

2022-02-07 05:52:00,476 yolo_tf2.utils.common.activate_gpu +325: INFO     [260] GPU activated
2022-02-07 05:52:00,477 yolo_tf2.utils.common.train +468: INFO     [260] Starting training ...
2022-02-07 05:52:04,293 yolo_tf2.utils.common.create_models +447: INFO     [260] Training and inference models created
2022-02-07 05:52:04,295 yolo_tf2.utils.common.wrapper +64: INFO     [260] create_models execution time: 3.8118433569999866 seconds
2022-02-07 05:52:04,301 yolo_tf2.utils.common.create_new_dataset +366: INFO     [260] Generating new dataset ...
2022-02-07 05:52:07,014 yolo_tf2.utils.common.adjust_non_voc_csv +184: INFO     [260] Adjustment from existing received 10107 labels containing 16 classes
2022-02-07 05:52:07,022 yolo_tf2.utils.common.adjust_non_voc_csv +187: INFO     [260] Added prefix to images: /content/yolo-data/images
Parsed labels:
Car               3153
Pedestrian        1418
Palm Tree         1379
Traffic Lights    1269
Street Sign       1109
Street Lamp        995
Road Block         363
Flag               124
Trash Can           90
Minivan             68
Fire Hydrant        52
Bus                 43
Pickup Truck        20
Bicycle             17
Delivery Truck       4
Motorcycle           3
Name: object_name, dtype: int64
2022-02-07 05:52:09,513 yolo_tf2.utils.common.save_fig +33: INFO     [260] Saved figure /content/output/plots/Relative width and height for 10107 boxes..png
/usr/local/lib/python3.7/dist-packages/yolo_tf2/utils/dataset_handlers.py:209: VisibleDeprecationWarning: Creating an ndarray from ragged nested sequences (which is a list-or-tuple of lists-or-tuples-or ndarrays with different lengths or shapes) is deprecated. If you meant to do this, you must specify 'dtype=object' when creating the ndarray
  groups = np.array(data.groupby('image_path'))
Processing beverly_hills_train.tfrecord
Building example: 406/411 ... Beverly_hills184.jpg 99% completed2022-02-07 05:52:12,922 yolo_tf2.utils.common.save_tfr +227: INFO     [260] Saved training TFRecord: /content/data/tfrecords/beverly_hills_train.tfrecord
Building example: 411/411 ... Beverly_hills365.jpg 100% completed
Processing beverly_hills_test.tfrecord
Building example: 31/46 ... Beverly_hills335.jpg 67% completed2022-02-07 05:52:13,175 yolo_tf2.utils.common.save_tfr +229: INFO     [260] Saved validation TFRecord: /content/data/tfrecords/beverly_hills_test.tfrecord
2022-02-07 05:52:13,271 yolo_tf2.utils.common.read_tfr +263: INFO     [260] Read TFRecord: /content/data/tfrecords/beverly_hills_train.tfrecord
Building example: 46/46 ... Beverly_hills186.jpg 100% completed
2022-02-07 05:52:18,892 yolo_tf2.utils.common.read_tfr +263: INFO     [260] Read TFRecord: /content/data/tfrecords/beverly_hills_test.tfrecord
/usr/local/lib/python3.7/dist-packages/keras/engine/functional.py:1410: CustomMaskWarning: Custom mask layers require a config and must override get_config. When loading, the custom mask layer must be passed to the custom_objects argument.
  layer_config = serialize_layer_fn(layer)
2022-02-07 05:52:50.575910: W tensorflow/python/util/util.cc:368] Sets are not currently considered sequences, but this may change in the future, so consider avoiding using them.
INFO:tensorflow:Assets written to: ram://eefa3127-ad7d-4445-a186-75fd8f0b81e1/assets
/usr/local/lib/python3.7/dist-packages/keras/engine/functional.py:1410: CustomMaskWarning: Custom mask layers require a config and must override get_config. When loading, the custom mask layer must be passed to the custom_objects argument.
  layer_config = serialize_layer_fn(layer)
INFO:tensorflow:Assets written to: ram://cbe6d5a4-5322-494b-ba91-3fd34131cdd9/assets
/usr/local/lib/python3.7/dist-packages/keras/engine/functional.py:1410: CustomMaskWarning: Custom mask layers require a config and must override get_config. When loading, the custom mask layer must be passed to the custom_objects argument.
  layer_config = serialize_layer_fn(layer)
INFO:tensorflow:Assets written to: ram://f15f3f25-9adb-4eb0-aa0d-83fa874bc74e/assets
/usr/local/lib/python3.7/dist-packages/keras/engine/functional.py:1410: CustomMaskWarning: Custom mask layers require a config and must override get_config. When loading, the custom mask layer must be passed to the custom_objects argument.
  layer_config = serialize_layer_fn(layer)
INFO:tensorflow:Assets written to: ram://86dd6f5f-4416-4465-99c0-928fd88e8a93/assets
/usr/local/lib/python3.7/dist-packages/keras/engine/functional.py:1410: CustomMaskWarning: Custom mask layers require a config and must override get_config. When loading, the custom mask layer must be passed to the custom_objects argument.
  layer_config = serialize_layer_fn(layer)
INFO:tensorflow:Assets written to: ram://ca08220f-cabc-4017-96d3-383557342388/assets
/usr/local/lib/python3.7/dist-packages/keras/engine/functional.py:1410: CustomMaskWarning: Custom mask layers require a config and must override get_config. When loading, the custom mask layer must be passed to the custom_objects argument.
  layer_config = serialize_layer_fn(layer)
INFO:tensorflow:Assets written to: ram://0f634207-e822-4d6c-a805-3cfeab37532f/assets
/usr/local/lib/python3.7/dist-packages/keras/engine/functional.py:1410: CustomMaskWarning: Custom mask layers require a config and must override get_config. When loading, the custom mask layer must be passed to the custom_objects argument.
  layer_config = serialize_layer_fn(layer)
INFO:tensorflow:Assets written to: ram://a971d021-3da4-402a-a004-4ae4aa67148a/assets
/usr/local/lib/python3.7/dist-packages/keras/engine/functional.py:1410: CustomMaskWarning: Custom mask layers require a config and must override get_config. When loading, the custom mask layer must be passed to the custom_objects argument.
  layer_config = serialize_layer_fn(layer)
INFO:tensorflow:Assets written to: ram://31d72fdf-1ce6-4131-a7e6-f6444747e9c9/assets
/usr/local/lib/python3.7/dist-packages/keras/engine/functional.py:1410: CustomMaskWarning: Custom mask layers require a config and must override get_config. When loading, the custom mask layer must be passed to the custom_objects argument.
  layer_config = serialize_layer_fn(layer)
INFO:tensorflow:Assets written to: ram://dac323b6-591a-481c-bbe6-85bb82bef38c/assets
/usr/local/lib/python3.7/dist-packages/keras/engine/functional.py:1410: CustomMaskWarning: Custom mask layers require a config and must override get_config. When loading, the custom mask layer must be passed to the custom_objects argument.
  layer_config = serialize_layer_fn(layer)
INFO:tensorflow:Assets written to: ram://99b029f7-11d1-40f2-b459-fd1d8dca5ba1/assets
/usr/local/lib/python3.7/dist-packages/keras/engine/functional.py:1410: CustomMaskWarning: Custom mask layers require a config and must override get_config. When loading, the custom mask layer must be passed to the custom_objects argument.
  layer_config = serialize_layer_fn(layer)
INFO:tensorflow:Assets written to: ram://210489fb-0895-4769-8be3-effd01d92695/assets
/usr/local/lib/python3.7/dist-packages/keras/engine/functional.py:1410: CustomMaskWarning: Custom mask layers require a config and must override get_config. When loading, the custom mask layer must be passed to the custom_objects argument.
  layer_config = serialize_layer_fn(layer)

Here's the output without the problem (tensorflow 2.5):

2022-02-07 06:09:53.125735: I tensorflow/stream_executor/platform/default/dso_loader.cc:53] Successfully opened dynamic library libcudart.so.11.0
2022-02-07 06:09:55.370728: I tensorflow/stream_executor/platform/default/dso_loader.cc:53] Successfully opened dynamic library libcuda.so.1
2022-02-07 06:09:55,387 yolo_tf2.utils.common.train +468: INFO     [269] Starting training ...
2022-02-07 06:09:55.387211: E tensorflow/stream_executor/cuda/cuda_driver.cc:328] failed call to cuInit: CUDA_ERROR_NO_DEVICE: no CUDA-capable device is detected
2022-02-07 06:09:55.387252: I tensorflow/stream_executor/cuda/cuda_diagnostics.cc:156] kernel driver does not appear to be running on this host (de0312867ce7): /proc/driver/nvidia/version does not exist
2022-02-07 06:09:55.427963: I tensorflow/core/platform/cpu_feature_guard.cc:142] This TensorFlow binary is optimized with oneAPI Deep Neural Network Library (oneDNN) to use the following CPU instructions in performance-critical operations:  AVX2 FMA
To enable them in other operations, rebuild TensorFlow with the appropriate compiler flags.
2022-02-07 06:10:00,078 yolo_tf2.utils.common.create_models +447: INFO     [269] Training and inference models created
2022-02-07 06:10:00,080 yolo_tf2.utils.common.wrapper +64: INFO     [269] create_models execution time: 4.689235652999997 seconds
2022-02-07 06:10:00,081 yolo_tf2.utils.common.create_new_dataset +366: INFO     [269] Generating new dataset ...
2022-02-07 06:10:02,572 yolo_tf2.utils.common.adjust_non_voc_csv +184: INFO     [269] Adjustment from existing received 10107 labels containing 16 classes
2022-02-07 06:10:02,574 yolo_tf2.utils.common.adjust_non_voc_csv +187: INFO     [269] Added prefix to images: /content/yolo-data/images
Parsed labels:
Car               3153
Pedestrian        1418
Palm Tree         1379
Traffic Lights    1269
Street Sign       1109
Street Lamp        995
Road Block         363
Flag               124
Trash Can           90
Minivan             68
Fire Hydrant        52
Bus                 43
Pickup Truck        20
Bicycle             17
Delivery Truck       4
Motorcycle           3
Name: object_name, dtype: int64
2022-02-07 06:10:04,900 yolo_tf2.utils.common.save_fig +33: INFO     [269] Saved figure /content/output/plots/Relative width and height for 10107 boxes..png
/usr/local/lib/python3.7/dist-packages/yolo_tf2/utils/dataset_handlers.py:209: VisibleDeprecationWarning: Creating an ndarray from ragged nested sequences (which is a list-or-tuple of lists-or-tuples-or ndarrays with different lengths or shapes) is deprecated. If you meant to do this, you must specify 'dtype=object' when creating the ndarray
  groups = np.array(data.groupby('image_path'))
Processing beverly_hills_train.tfrecord
Building example: 392/411 ... Beverly_hills294.jpg 95% completed2022-02-07 06:10:10,341 yolo_tf2.utils.common.save_tfr +227: INFO     [269] Saved training TFRecord: /content/data/tfrecords/beverly_hills_train.tfrecord
Building example: 411/411 ... Beverly_hills94.jpg 100% completed
Processing beverly_hills_test.tfrecord
Building example: 25/46 ... Beverly_hills334.jpg 54% completed2022-02-07 06:10:10,730 yolo_tf2.utils.common.save_tfr +229: INFO     [269] Saved validation TFRecord: /content/data/tfrecords/beverly_hills_test.tfrecord
Building example: 46/46 ... Beverly_hills251.jpg 100% completed
2022-02-07 06:10:10,843 yolo_tf2.utils.common.read_tfr +263: INFO     [269] Read TFRecord: /content/data/tfrecords/beverly_hills_train.tfrecord
2022-02-07 06:10:15,264 yolo_tf2.utils.common.read_tfr +263: INFO     [269] Read TFRecord: /content/data/tfrecords/beverly_hills_test.tfrecord
2022-02-07 06:10:15.676352: I tensorflow/core/profiler/lib/profiler_session.cc:126] Profiler session initializing.
2022-02-07 06:10:15.676423: I tensorflow/core/profiler/lib/profiler_session.cc:141] Profiler session started.
2022-02-07 06:10:15.701051: I tensorflow/core/profiler/lib/profiler_session.cc:159] Profiler session tear down.
/usr/local/lib/python3.7/dist-packages/tensorflow/python/keras/utils/generic_utils.py:497: CustomMaskWarning: Custom mask layers require a config and must override get_config. When loading, the custom mask layer must be passed to the custom_objects argument.
  category=CustomMaskWarning)
2022-02-07 06:10:17.064324: I tensorflow/compiler/mlir/mlir_graph_optimization_pass.cc:176] None of the MLIR Optimization Passes are enabled (registered 2)
2022-02-07 06:10:17.081408: I tensorflow/core/platform/profile_utils/cpu_utils.cc:114] CPU Frequency: 2199995000 Hz
Epoch 1/100
      1/Unknown - 40s 40s/step - loss: 7333.2617 - layer_205_lambda_loss: 403.8862 - layer_230_lambda_loss: 1509.9465 - layer_255_lambda_loss: 5407.68902022-02-07 06:10:59.974130: I tensorflow/core/profiler/lib/profiler_session.cc:126] Profiler session initializing.
2022-02-07 06:10:59.974196: I tensorflow/core/profiler/lib/profiler_session.cc:141] Profiler session started.
      2/Unknown - 50s 11s/step - loss: 7819.7124 - layer_205_lambda_loss: 697.4546 - layer_230_lambda_loss: 1647.7856 - layer_255_lambda_loss: 5462.71582022-02-07 06:11:10.059899: I tensorflow/core/profiler/lib/profiler_session.cc:66] Profiler session collecting data.
2022-02-07 06:11:10.088821: I tensorflow/core/profiler/lib/profiler_session.cc:159] Profiler session tear down.
2022-02-07 06:11:10.133747: I tensorflow/core/profiler/rpc/client/save_profile.cc:137] Creating directory: /content/data/tfrecords/train/plugins/profile/2022_02_07_06_11_10
2022-02-07 06:11:10.157875: I tensorflow/core/profiler/rpc/client/save_profile.cc:143] Dumped gzipped tool data for trace.json.gz to /content/data/tfrecords/train/plugins/profile/2022_02_07_06_11_10/de0312867ce7.trace.json.gz
2022-02-07 06:11:10.189438: I tensorflow/core/profiler/rpc/client/save_profile.cc:137] Creating directory: /content/data/tfrecords/train/plugins/profile/2022_02_07_06_11_10
2022-02-07 06:11:10.189678: I tensorflow/core/profiler/rpc/client/save_profile.cc:143] Dumped gzipped tool data for memory_profile.json.gz to /content/data/tfrecords/train/plugins/profile/2022_02_07_06_11_10/de0312867ce7.memory_profile.json.gz
2022-02-07 06:11:10.192796: I tensorflow/core/profiler/rpc/client/capture_profile.cc:251] Creating directory: /content/data/tfrecords/train/plugins/profile/2022_02_07_06_11_10Dumped tool data for xplane.pb to /content/data/tfrecords/train/plugins/profile/2022_02_07_06_11_10/de0312867ce7.xplane.pb
Dumped tool data for overview_page.pb to /content/data/tfrecords/train/plugins/profile/2022_02_07_06_11_10/de0312867ce7.overview_page.pb
Dumped tool data for input_pipeline.pb to /content/data/tfrecords/train/plugins/profile/2022_02_07_06_11_10/de0312867ce7.input_pipeline.pb
Dumped tool data for tensorflow_stats.pb to /content/data/tfrecords/train/plugins/profile/2022_02_07_06_11_10/de0312867ce7.tensorflow_stats.pb
Dumped tool data for kernel_stats.pb to /content/data/tfrecords/train/plugins/profile/2022_02_07_06_11_10/de0312867ce7.kernel_stats.pb

     15/Unknown - 181s 10s/step - loss: 3493.4009 - layer_205_lambda_loss: 232.7220 - layer_230_lambda_loss: 629.6332 - layer_255_lambda_loss: 2618.9722

I also tried (same results):

  • python versions: 3.8, 3.9, 3.10
  • Ubuntu 18, and OSX
  • tensorflow versions: 2.7, 2.8, 2.9.0-dev20220203
1 Answers

I can confirm this same error is still encountered in tensorflow 2.8 even with the use of the latest nightly version of 2.8.0 (2.8.0-dev20211222). This was encountered with several tf.keras models (where some work fine in tensorflow 2.2 version but simply stop training after few epochs in 2.8 version). Even with very small datasets. Decreasing the batch size improves a bit the training by delaying the error explosion to later epochs, but it doesn't solve the issue.

I have tried to use os.environ["TF_GPU_ALLOCATOR"]="cuda_malloc_async" but in vain.

Here is an example of an error output when training (memory usage keeps increasing during training, until the 60th epoch where an OOM memory burst problem happens).

2022-03-11 16:33:20.651586: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1525] Created device /job:localhost/replica:0/task:0/device:GPU:0 with 11433 MB memory: -> device: 0, name: NVIDIA TITAN X (Pascal), pci bus id: 0000:02:00.0, compute capability: 6.1 2022-03-11 16:33:35.198049: I tensorflow/stream_executor/cuda/cuda_dnn.cc:368] Loaded cuDNN version 8101 2022-03-11 16:42:26.891801: W tensorflow/core/common_runtime/bfc_allocator.cc:462] Allocator (GPU_0_bfc) ran out of memory trying to allocate 98.00MiB (rounded to 102760448)requested by op gradient_tape/model_2/conv_3_2/Conv2D/Conv2DBackpropInput If the cause is memory fragmentation maybe the environment variable 'TF_GPU_ALLOCATOR=cuda_malloc_async' will improve the situation. Current allocation summary follows. Current allocation summary follows. 2022-03-11 16:42:26.891975: I tensorflow/core/common_runtime/bfc_allocator.cc:1010] BFCAllocator dump for GPU_0_bfc 2022-03-11 16:42:26.892028: I tensorflow/core/common_runtime/bfc_allocator.cc:1017] Bin (256): Total Chunks: 179, Chunks in use: 173. 44.8KiB allocated for chunks. 43.2KiB in use in bin. 16.7KiB client-requested in use in bin. 2022-03-11 16:42:26.892060: I tensorflow/core/common_runtime/bfc_allocator.cc:1017] Bin (512): Total Chunks: 58, Chunks in use: 52. 29.2KiB allocated for chunks. 26.2KiB in use in bin. 26.0KiB client-requested in use in bin. 2022-03-11 16:42:26.892093: I tensorflow/core/common_runtime/bfc_allocator.cc:1017] Bin (1024): Total Chunks: 203, Chunks in use: 202. 212.5KiB allocated for chunks. 211.5KiB in use in bin. 202.0KiB client-requested in use in bin. 2022-03-11 16:42:26.892639: I tensorflow/core/common_runtime/bfc_allocator.cc:1017] Bin (2048): Total Chunks: 1, Chunks in use: 0. 2.0KiB allocated for chunks. 0B in use in bin. 0B client-requested in use in bin. 2022-03-11 16:42:26.892864: I tensorflow/core/common_runtime/bfc_allocator.cc:1017] Bin (4096): Total Chunks: 0, Chunks in use: 0. 0B allocated for chunks. 0B in use in bin. 0B client-requested in use in bin. 2022-03-11 16:42:26.892915: I tensorflow/core/common_runtime/bfc_allocator.cc:1017] Bin (8192): Total Chunks: 22, Chunks in use: 22. 179.0KiB allocated for chunks. 179.0KiB in use in bin. 176.0KiB client-requested in use in bin. 2022-03-11 16:42:26.892954: I tensorflow/core/common_runtime/bfc_allocator.cc:1017] Bin (16384): Total Chunks: 21, Chunks in use: 18. 518.8KiB allocated for chunks. 446.2KiB in use in bin. 442.0KiB client-requested in use in bin. 2022-03-11 16:42:26.892989: I tensorflow/core/common_runtime/bfc_allocator.cc:1017] Bin (32768): Total Chunks: 4, Chunks in use: 3. 184.0KiB allocated for chunks. 134.0KiB in use in bin. 81.0KiB client-requested in use in bin. 2022-03-11 16:42:26.893018: I tensorflow/core/common_runtime/bfc_allocator.cc:1017] Bin (65536): Total Chunks: 0, Chunks in use: 0. 0B allocated for chunks. 0B in use in bin. 0B client-requested in use in bin. 2022-03-11 16:42:26.893050: I tensorflow/core/common_runtime/bfc_allocator.cc:1017] Bin (131072): Total Chunks: 6, Chunks in use: 5. 948.5KiB allocated for chunks. 804.5KiB in use in bin. 720.0KiB client-requested in use in bin. 2022-03-11 16:42:26.893079: I tensorflow/core/common_runtime/bfc_allocator.cc:1017] Bin (262144): Total Chunks: 5, Chunks in use: 5. 1.58MiB allocated for chunks. 1.58MiB in use in bin. 1.41MiB client-requested in use in bin. 2022-03-11 16:42:26.893297: I tensorflow/core/common_runtime/bfc_allocator.cc:1017] Bin (524288): Total Chunks: 18, Chunks in use: 14. 10.70MiB allocated for chunks. 8.12MiB in use in bin. 7.88MiB client-requested in use in bin. 2022-03-11 16:42:26.893332: I tensorflow/core/common_runtime/bfc_allocator.cc:1017] Bin (1048576): Total Chunks: 8, Chunks in use: 5. 9.51MiB allocated for chunks. 6.06MiB in use in bin. 5.65MiB client-requested in use in bin. 2022-03-11 16:42:26.893367: I tensorflow/core/common_runtime/bfc_allocator.cc:1017] Bin (2097152): Total Chunks: 19, Chunks in use: 19. 44.25MiB allocated for chunks. 44.25MiB in use in bin. 42.75MiB client-requested in use in bin. 2022-03-11 16:42:26.893402: I tensorflow/core/common_runtime/bfc_allocator.cc:1017] Bin (4194304): Total Chunks: 12, Chunks in use: 10. 68.09MiB allocated for chunks. 54.09MiB in use in bin. 53.22MiB client-requested in use in bin. 2022-03-11 16:42:26.893431: I tensorflow/core/common_runtime/bfc_allocator.cc:1017] Bin (8388608): Total Chunks: 0, Chunks in use: 0. 0B allocated for chunks. 0B in use in bin. 0B client-requested in use in bin. 2022-03-11 16:42:26.893464: I tensorflow/core/common_runtime/bfc_allocator.cc:1017] Bin (16777216): Total Chunks: 15, Chunks in use: 14. 360.07MiB allocated for chunks. 335.83MiB in use in bin. 311.38MiB client-requested in use in bin. 2022-03-11 16:42:26.893497: I tensorflow/core/common_runtime/bfc_allocator.cc:1017] Bin (33554432): Total Chunks: 13, Chunks in use: 12. 615.82MiB allocated for chunks. 578.18MiB in use in bin. 563.50MiB client-requested in use in bin. 2022-03-11 16:42:26.893709: I tensorflow/core/common_runtime/bfc_allocator.cc:1017] Bin (67108864): Total Chunks: 78, Chunks in use: 78. 6.23GiB allocated for chunks. 6.23GiB in use in bin. 5.94GiB client-requested in use in bin. 2022-03-11 16:42:26.893742: I tensorflow/core/common_runtime/bfc_allocator.cc:1017] Bin (134217728): Total Chunks: 7, Chunks in use: 7. 1.29GiB allocated for chunks. 1.29GiB in use in bin. 1.24GiB client-requested in use in bin. 2022-03-11 16:42:26.893774: I tensorflow/core/common_runtime/bfc_allocator.cc:1017] Bin (268435456): Total Chunks: 6, Chunks in use: 6. 2.56GiB allocated for chunks. 2.56GiB in use in bin. 2.45GiB client-requested in use in bin. 2022-03-11 16:42:26.895086: I tensorflow/core/common_runtime/bfc_allocator.cc:1033] Bin for 98.00MiB was 64.00MiB, Chunk State: 2022-03-11 16:42:26.895129: I tensorflow/core/common_runtime/bfc_allocator.cc:1046] Next region of size 11988434944 2022-03-11 16:42:26.895161: I tensorflow/core/common_runtime/bfc_allocator.cc:1066] InUse at 7fd8cc000000 of size 1280 next 1 2022-03-11 16:42:26.895380: I tensorflow/core/common_runtime/bfc_allocator.cc:1066] InUse at 7fd8cc000500 of size 256 next 2 2022-03-11 16:42:26.895411: I tensorflow/core/common_runtime/bfc_allocator.cc:1066] InUse at 7fd8cc000600 of size 256 next 3 2022-03-11 16:42:26.895433: I tensorflow/core/common_runtime/bfc_allocator.cc:1066] InUse at 7fd8cc000700 of size 256 next 4 2022-03-11 16:42:26.895620: I tensorflow/core/common_runtime/bfc_allocator.cc:1066] InUse at 7fd8cc000800 of size 256 next 5 2022-03-11 16:42:26.895656: I tensorflow/core/common_runtime/bfc_allocator.cc:1066] InUse at 7fd8cc000900 of size 256 next 6 2022-03-11 16:42:26.895680: I tensorflow/core/common_runtime/bfc_allocator.cc:1066] InUse at 7fd8cc000a00 of size 256 next 7 2022-03-11 16:42:26.895701: I tensorflow/core/common_runtime/bfc_allocator.cc:1066] InUse at 7fd8cc000b00 of size 256 next 8 2022-03-11 16:42:26.895721: I tensorflow/core/common_runtime/bfc_allocator.cc:1066] InUse at 7fd8cc000c00 of size 256 next 11 2022-03-11 16:42:26.895742: I tensorflow/core/common_runtime/bfc_allocator.cc:1066] InUse at ... tensorflow/core/common_runtime/bfc_allocator.cc:1074] 1 Chunks of size 77690368 totalling 74.09MiB 2022-03-11 16:42:26.914528: I tensorflow/core/common_runtime/bfc_allocator.cc:1074] 1 Chunks of size 82451456 totalling 78.63MiB 2022-03-11 16:42:26.914551: I tensorflow/core/common_runtime/bfc_allocator.cc:1074] 1 Chunks of size 82456576 totalling 78.64MiB 2022-03-11 16:42:26.914573: I tensorflow/core/common_runtime/bfc_allocator.cc:1074] 1 Chunks of size 87380992 totalling 83.33MiB 2022-03-11 16:42:26.914594: I tensorflow/core/common_runtime/bfc_allocator.cc:1074] 1 Chunks of size 89632256 totalling 85.48MiB 2022-03-11 16:42:26.914617: I tensorflow/core/common_runtime/bfc_allocator.cc:1074] 1 Chunks of size 90040832 totalling 85.87MiB 2022-03-11 16:42:26.914639: I tensorflow/core/common_runtime/bfc_allocator.cc:1074] 1 Chunks of size 95168768 totalling 90.76MiB 2022-03-11 16:42:26.914660: I tensorflow/core/common_runtime/bfc_allocator.cc:1074] 1 Chunks of size 96234496 totalling 91.78MiB 2022-03-11 16:42:26.914682: I tensorflow/core/common_runtime/bfc_allocator.cc:1074] 1 Chunks of size 99255296 totalling 94.66MiB 2022-03-11 16:42:26.914705: I tensorflow/core/common_runtime/bfc_allocator.cc:1074] 1 Chunks of size 99264512 totalling 94.67MiB 2022-03-11 16:42:26.914726: I tensorflow/core/common_runtime/bfc_allocator.cc:1074] 13 Chunks of size 102760448 totalling 1.24GiB 2022-03-11 16:42:26.914747: I tensorflow/core/common_runtime/bfc_allocator.cc:1074] 1 Chunks of size 104173312 totalling 99.35MiB 2022-03-11 16:42:26.914770: I tensorflow/core/common_runtime/bfc_allocator.cc:1074] 1 Chunks of size 105158656 totalling 100.29MiB 2022-03-11 16:42:26.914792: I tensorflow/core/common_runtime/bfc_allocator.cc:1074] 1 Chunks of size 108818432 totalling 103.78MiB 2022-03-11 16:42:26.914812: I tensorflow/core/common_runtime/bfc_allocator.cc:1074] 1 Chunks of size 110231552 totalling 105.12MiB 2022-03-11 16:42:26.914832: I tensorflow/core/common_runtime/bfc_allocator.cc:1074] 1 Chunks of size 119567360 totalling 114.03MiB 2022-03-11 16:42:26.914877: I tensorflow/core/common_runtime/bfc_allocator.cc:1074] 1 Chunks of size 127390208 totalling 121.49MiB 2022-03-11 16:42:26.914900: I tensorflow/core/common_runtime/bfc_allocator.cc:1074] 1 Chunks of size 154140672 totalling 147.00MiB 2022-03-11 16:42:26.914921: I tensorflow/core/common_runtime/bfc_allocator.cc:1074] 6 Chunks of size 205520896 totalling 1.15GiB 2022-03-11 16:42:26.914943: I tensorflow/core/common_runtime/bfc_allocator.cc:1074] 5 Chunks of size 411041792 totalling 1.91GiB 2022-03-11 16:42:26.914965: I tensorflow/core/common_runtime/bfc_allocator.cc:1074] 1 Chunks of size 692452608 totalling 660.37MiB 2022-03-11 16:42:26.914986: I tensorflow/core/common_runtime/bfc_allocator.cc:1078] Sum Total of in-use chunks: 11.08GiB 2022-03-11 16:42:26.915006: I tensorflow/core/common_runtime/bfc_allocator.cc:1080] total_region_allocated_bytes_: 11988434944 memory_limit_: 11988434944 available bytes: 0 curr_region_allocation_bytes_: 23976869888 2022-03-11 16:42:26.915035: I tensorflow/core/common_runtime/bfc_allocator.cc:1086] Stats: Limit:
11988434944 InUse: 11902258688 MaxInUse:
11902527488 NumAllocs: 1785814 MaxAllocSize:
1866465280 Reserved: 0 PeakReserved:
0 LargestFreeBlock: 0

2022-03-11 16:42:26.915131: W tensorflow/core/common_runtime/bfc_allocator.cc:474] ***************************************************************************************************x 2022-03-11 16:42:26.916684: W tensorflow/core/framework/op_kernel.cc:1745] OP_REQUIRES failed at conv_grad_input_ops.cc:408 : RESOURCE_EXHAUSTED: OOM when allocating tensor with shape[128,256,28,28] and type float on /job:localhost/replica:0/task:0/device:GPU:0 by allocator GPU_0_bfc

Is there any clue on how to prevent this and if this error has or is been addressed ?

I would be happy to provide further details if needed. Thanks.

Related