Model training stops at random epochs. First few epochs run to completion. Apple M1 Max Mac Studio

Viewed 93

My code keeps getting stuck at random epochs every time I run it. I just got tensor flow on my M1 Max and I need to know if it is an issue with that. I can post the YAML file I used to set up the mini forge environment as well. Most of the answers online ask me to uninstall tensorflow-metal which I do not want to do. Thank you for the help in advance!

Mainly this is the error that I get after the first time it gets stuck, i.e. it only has this line from the second run onwards.

Could not identify NUMA node of platform GPU ID 0, defaulting to 0. Your kernel may not have been built with NUMA support.

Here is the entire error message when it is stuck.

Metal device set to: Apple M1 Max

systemMemory: 64.00 GB
maxCacheSize: 24.00 GB

Epoch 1/100
2022-07-08 17:05:12.634579: I tensorflow/core/common_runtime/pluggable_device/pluggable_device_factory.cc:305] Could not identify NUMA node of platform GPU ID 0, defaulting to 0. Your kernel may not have been built with NUMA support.
2022-07-08 17:05:12.634708: I tensorflow/core/common_runtime/pluggable_device/pluggable_device_factory.cc:271] Created TensorFlow device (/job:localhost/replica:0/task:0/device:GPU:0 with 0 MB memory) -> physical PluggableDevice (device: 0, name: METAL, pci bus id: <undefined>)
2022-07-08 17:05:12.779355: W tensorflow/core/platform/profile_utils/cpu_utils.cc:128] Failed to get CPU frequency: 0 Hz
2022-07-08 17:05:13.161749: I tensorflow/core/grappler/optimizers/custom_graph_optimizer_registry.cc:113] Plugin optimizer for device_type GPU is enabled.
Output exceeds the size limit. Open the full output data in a text editor
355/355 [==============================] - 3s 8ms/step - loss: 8771.2441 - mse: 8771.2441
Epoch 2/100
355/355 [==============================] - 3s 8ms/step - loss: 5875.3423 - mse: 5875.3423
Epoch 3/100
355/355 [==============================] - 3s 8ms/step - loss: 3721.3096 - mse: 3721.3096
Epoch 4/100
355/355 [==============================] - 3s 8ms/step - loss: 1454.0521 - mse: 1454.0521
Epoch 5/100
355/355 [==============================] - 3s 8ms/step - loss: 1102.7731 - mse: 1102.7731
Epoch 6/100
355/355 [==============================] - 3s 8ms/step - loss: 991.1359 - mse: 991.1359
Epoch 7/100
355/355 [==============================] - 3s 8ms/step - loss: 959.4382 - mse: 959.4382
Epoch 8/100
355/355 [==============================] - 3s 8ms/step - loss: 905.4555 - mse: 905.4555
Epoch 9/100
355/355 [==============================] - 3s 8ms/step - loss: 871.3311 - mse: 871.3311
Epoch 10/100
355/355 [==============================] - 3s 8ms/step - loss: 799.5955 - mse: 799.5955
Epoch 11/100
355/355 [==============================] - 3s 8ms/step - loss: 786.7227 - mse: 786.7227
Epoch 12/100
355/355 [==============================] - 3s 8ms/step - loss: 789.5096 - mse: 789.5096
Epoch 13/100
355/355 [==============================] - 3s 8ms/step - loss: 753.4902 - mse: 753.4902
...
355/355 [==============================] - 3s 8ms/step - loss: 615.7115 - mse: 615.7115
Epoch 57/100
355/355 [==============================] - 133983s 378s/step - loss: 617.2227 - mse: 617.2227
Epoch 58/100
235/355 [==================>...........] - ETA: 14:53:06 - loss: 609.9565 - mse: 609.9565

And here is the code (just for the architecture) if that is useful:

X_train_tensor = tf.constant( X_train, dtype=tf.float32)
Y_train_tensor = tf.constant( Y_train, dtype=tf.float32)

X_test_tensor = tf.constant( X_test, dtype=tf.float32)
Y_test_tensor = tf.constant( Y_test, dtype=tf.float32)



model = tf.keras.Sequential()
model.add(tf.keras.layers.Dense(X_train.shape[1], activation='relu', input_shape=[X_train.shape[1]]))
model.add(tf.keras.layers.Dense(16, activation='relu'))
model.add(tf.keras.layers.Dropout(0.1))
model.add(tf.keras.layers.Dense(50, activation='relu'))
model.add(tf.keras.layers.Dropout(0.1))
model.add(tf.keras.layers.Dense(400, activation='relu'))
model.add(tf.keras.layers.Dropout(0.1))
model.add(tf.keras.layers.Dense(50, activation='relu'))
model.add(tf.keras.layers.Dropout(0.1))
model.add(tf.keras.layers.Dense(16, activation='relu'))
model.add(tf.keras.layers.Dense(X_train.shape[1], activation='relu'))
model.add(tf.keras.layers.Dropout(0.1))
model.add(tf.keras.layers.Dense(1))
#
model.compile(optimizer= 'adam', loss='mse', metrics=['mse'])
history = model.fit(X_train_tensor, Y_train_tensor, epochs=100)

TensorFlow-metal: 0.5.0 tensorflow-macos: 2.9.2 python: 3.9.13 keras: 2.9.0

0 Answers
Related