How to do multi GPU training with Keras?

Viewed 6138

I want my model to run on multiple GPU-sharing parameters but with different batches of data.

Can I do something like that with model.fit()? Is there any other alternative?

3 Answers

In multi-gpu model training is very convenient than ever. Check the following document regarding this: Multi-GPU and distributed training.


In essence, to do single-host, multi-device synchronous training with a model, you would use the tf.distribute.MirroredStrategy API. Here's how it works:

  • Instantiate a MirroredStrategy, optionally configuring which specific devices you want to use (by default the strategy will use all GPUs available).

  • Use the strategy object to open a scope, and within this scope, create all the Keras objects you need that contain variables. Typically, that means creating & compiling the model inside the distribution scope.

  • Train the model via fit() as usual.

Schematically, it looks like this:

# Create a MirroredStrategy.
strategy = tf.distribute.MirroredStrategy()
print('Number of devices: {}'.format(strategy.num_replicas_in_sync))

# Open a strategy scope.
with strategy.scope():
  # Everything that creates variables should be under the strategy scope.
  # In general this is only model construction & `compile()`.
  model = Model(...)
  model.compile(...)

# Train the model on all available devices.
model.fit(train_dataset, validation_data=val_dataset, ...)

# Test the model on all available devices.
model.evaluate(test_dataset)
Related