Assume there are two models to be used: X and Y. Data is sequentially passed through X and Y. Only the parameters of model Y need to be optimized with respect to a loss computed over output of model Y. Is the following snippet a correct implementation of this requirement. Few specific queries I need answers to:
- What does the
with torch.no_grad()exactly do? - As only the parameters of model Y are registered with the optimizer do we still need to freeze model X to be correct or is it only required to reduce the computational load?
- More generally I want an explanation on how the computation graph and backpropagation behaves in the presence of
with torch.no_grad()or when some layers are freezed by setting the correspondingrequires_gradparameter to False. - Also comment on whether we can have non-consecutive layers in the network frozen at once.
optimizer = AdamW(model_Y.parameters(), lr= , eps= , ...)
optimizer.zero_grad()
with torch.no_grad():
A = model_X(data)
B = model_Y(A)
loss = some_function(B)
loss.backward()
optimizer.step()