Given an HxW binary image (represented as a numpy 2d array) and D (integer), I would like to output another HxW 2d array in which each (i, j) index stores the number of 1 pixels in the binary image which are at most D rows or columns (basically, up to D pixels away in the L1 sense) from (i, j).
I can of course achieve this by convolving the binary image with a DxD all-ones square, but that seems rather slow using scipy.signal.convolve2d, for example. Also, my D can be rather large (e.g. 256 for an image of size 2600x1900). Any other suggestions?