Keras multi GPU using Tensorflow MirroredStrategy

Viewed 629
es = EarlyStopping(monitor='val_loss', mode='min', patience=100, restore_best_weights=True, verbose=0)
strategy = tf.distribute.MirroredStrategy(devices=['/gpu:0', '/gpu:1', '/gpu:2', '/gpu:3'])
with strategy.scope():
   model = RESNET()
history = model.fit(samples2Fit, validation_data=samples2Validate, epochs=args.epochs, callbacks=[es], verbose=0)

The RESNET() model is compiled as: model.compile(loss=tf.keras.losses.Huber(), optimizer=tf.keras.optimizers.Adam(epsilon=1e-08), metrics=[tf.keras.losses.Huber()]) and all other modules are also from tensorflow.keras.**

When I run this using 4 GPUs I get the following error: ValueError: Please use tf.keras.losses.Reduction.SUM or tf.keras.losses.Reduction.NONE for loss reduction when losses are used with tf.distribute.Strategy outside of the built-in training loops...

I am following the example given in https://keras.io/guides/distributed_training/ so what am I missing and why do I need to use these reductions? What is meant by outside of the built-in training loops?

1 Answers

try fitting within the strategy scope

with strategy.scope():
   model = RESNET()
   history = model.fit(samples2Fit, validation_data=samples2Validate, 
         epochs=args.epochs, callbacks=[es], verbose=0)

by default MirroredStrategy will use cross_device_ops with NcclAllReduce()

cross_device_ops: optional, a descedant of CrossDeviceOps. If this is not set, NcclAllReduce() will be used by default. One would customize this if NCCL isn't available or if a special implementation that exploits the particular hardware is available.

you might try different cross_device_ops options https://www.tensorflow.org/api_docs/python/tf/distribute/MirroredStrategy

Related