In the OpenAI Five paper it is mentioned that the "Gradients are additionally clipped per parameter to be within between ±5√v where v is the running estimate of the second moment of the (unclipped) gradient.". This is something I would like to implement in my project, but I am not sure how to do it neihter in theory nor in practice.
From wikipedia I found out that the "The second central moment is the variance. The positive square root of the variance is the standard deviation [...]". My best guess regarding the "running estimate" is that it is the Exponential Moving Average. The gradients of a network can be accessed as this comment suggests.
From these I would assume that √v is the Exponential Running Average of the standard dev. of the gradients and could be calculated via:
estimate = alpha * torch.std(list(param.grad for param in model.parameters())) + (1-alpha) * estimate
Is my theory correct? Is there a better way to do it? Thanks in advance.
Edit: fixed gradient gathering after Mr. For Example"s answer.
