I have a 3 dimensional numpy array of shape 20X10_000X15 and I would like to multiply it against a 1 dimensional numpy array of shape 1X15.
I would expect that the fastest way to do this is to use numpy vectorized methods and avoid for loops altogether.
In practice it seems that a hybrid approach of using a small amount of looping is faster than the purely vectorized operation.
For example:
N = 10_000 # Purely vectorized method: Multiply 20XNX15 array against a 1X15 array a = np.random.rand(20, N, 15) b = np.random.rand(15) %timeit np.multiply(b, a) # Multiply NX15 array against a 1X15 array 20 times a = np.random.rand(N, 15) b = np.random.rand(15) %timeit for i in range(20): np.multiply(b, a) >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>> 7.93 ms ± 57.2 µs per loop (mean ± std. dev. of 7 runs, 100 loops each) 5.43 ms ± 223 µs per loop (mean ± std. dev. of 7 runs, 100 loops each)The trend persists until I decrease the array size to below 20X3000X15, at which point the vectorized approach is faster
Does anyone know the reason for this? Does numpy have a diminishing return on vectorized operations on large arrays?
Is there a better way to do this?