Does tf.stop_gradient actually help save GPU memory. I'm asking since some intermediate outputs of layers behind the stop_gradient might not have to be stored (which would've otherwise been necessary for gradient computation).
Does tf.stop_gradient actually help save GPU memory. I'm asking since some intermediate outputs of layers behind the stop_gradient might not have to be stored (which would've otherwise been necessary for gradient computation).
Yes it saves GPU Memory. From the documentation of tf.stop_gradient, it states When building ops to compute gradients, this op prevents the contribution of its inputs to be taken into account.
One more explanation about tf.stop_gradient(): It is an operation that acts as the identity function in the forward direction, but stops the accumulated gradient from flowing through that operator in the backward direction. It does not prevent backpropagation altogether, but instead prevents an individual tensor from contributing to the gradients that are computed for an expression.
It means that the Weights included in the Operation which is passed to tf.stop_gradient are not updated during Back Propagation.
Normally, during Back Propagation, Weights are updated using the Formula,
W(i+1) = W(i) - Learning Rate * Gradient
When we use tf.stop_gradient, the Weights which are passed as Inputs to this Op are maintained as Constants. Since the Computations during Back Propagation are reduced, Memory Consumption also decreases.
For more information regarding tf.stop_gradient, please refer Tensorflow documentation, Abhishek's Stack Overflow Answer, mrry's Stack Overflow Answer and this article.
Hope this helps. Happy Learning!