Question summary: How is the dimensionality of inputs and outputs handled in the backward pass of custom functions?
According to the manual, the basic structure of custom functions is the following:
class MyFunc(torch.autograd.Function):
@staticmethod
def forward(ctx, input): # f(x) = e^x
result = input.exp()
ctx.save_for_backward(result)
return result
@staticmethod
def backward(ctx, grad_output): # df(x) = e^x
result, = ctx.saved_tensors
return grad_output * result
For a single input and output dimension, this is perfectly fine and works like a charm. But for higher dimensions the backward pass becomes confusing. Apparently, PyTorch only accepts a result of backward that has the same dimensionality as the result of forward (for the same input). Returning a wrong shape yields a RuntimeError: Function MyFunc returned an invalid gradient at index 0 - got [*] but expected shape compatible with [*]. So I am wondering: What does backward actually compute?
Its not a Jacobian? For example, when I have a function f(x) = ( f_1(x_1, ... , x_n), ... , f_k(x_1, ... , x_n) ) with n inputs and k outputs, I would expect that a gradient calculation would yield a Jacobian matrix of dimension k*n. However, the PyTorch implementation expects just a vector of dimension n. So what does the backward result actually mean, it can't be the Jacobian?
And it does not handle batches? Moreover, what if I would like to push a batch of input vectors through this function, e.g. an input of dimension b*n with batch size b. Then, instead of something like b*k*n the gradient is expected to also have the shape b*n. Is it even intended to consider the processing of batches with custom functions?
None of these questions seems to be addressed in the manual and the provided examples are very simple, which does not help at all. Maybe there are formulas hidden somewhere that explain the background of the provided Function interface in more detail, but I haven't found them yet.
