Accumulate gradients with distributed strategy in Tensorflow 2

Viewed 238

I have implemented a distributed strategy to train my model on multiple GPUs.

strategy = tf.distribute.MirroredStrategy(devices=devices[:FLAGS.n_gpus])
strategy.run(fn=self.train_step, args=(model, data))

My model now got more complex and bigger and I had to reduce the batch size to fit it onto the GPUs. The gradient is quite noisy now and I want to increase the batch size again by accumulating gradients.

Now my question is: is this even possible when using a mirrored strategy? I know that loss and gradients are combined across the replicas anyway, so is there a way to sum them across the replicas AND e.g. a loop running over the batches? I tried the straight-forward thing and returned the per replica calculated gradients to add and apply them outside the strategy.run() like that:

for b in batches:
    per_replica_gradients = strategy.run(fn=self.train_step, args=(model, data))
    total_gradient += per_replica_gradients
optimizer.apply_gradients(zip(total_gradient, model.trainable_variables)

but Tensorflow tells me that this is not possible and the gradients have to be applied withing the strategy.run(). This also makes sense to me but I wonder whether there is a possibility to accumulate gradients AND use a mirrored strategy?

1 Answers

You could use tf.distribute.ReplicaContext.all_reduce: This differs from Strategy.reduce in that it is for replica context and does not copy the results to the host device. all_reduce should be typically used for reductions inside the training step such as gradients.

More details can be found in the document here.

Related