Pytorch 2D convolution is somewhat slow

Viewed 339

I am trying to replace a single 2D convolution layer with a relatively large kernel, with several 2D-Conv layers having much smaller kernels. Theoretically, the replacement should work much faster (in respect of the number of operations) but actually it does not.

Suppose my kernel is of size OxCxHxW (output channels, input channels, height, width, respectfully), then in theory, the computational cost is OxCxHxW per input pixel. However, it seems that larger kernels are computed much faster.

For example:

  1. a kernel of size 32x32x5x5 takes 46ms to run (baseline large kernel)
  2. a smaller kernel of size 32x32x1x1 takes 8 ms (supposed to be x25 faster than #1, about 2ms)
  3. a group-convolution with a kernel size of 32x1x5x5 takes about 9 ms, while the reduction in kernel size is much more significant (supposed to be x32 faster than #1, about 1.5ms).

I have also tried TensorFlow and got similar results.

I understand that larger kernels probably utilize the CPU (or GPU) better and that there is a function's call overhead, but the time reduction going from large to the small kernel is not as significant as I expected.

A colab notebook can be found here.

Can anyone suggest an idea to accelerate those small convolutions?

0 Answers
Related