numpy vectorized operations slower than looping + vectorization?

Viewed 195
  • I have a 3 dimensional numpy array of shape 20X10_000X15 and I would like to multiply it against a 1 dimensional numpy array of shape 1X15.

  • I would expect that the fastest way to do this is to use numpy vectorized methods and avoid for loops altogether.

  • In practice it seems that a hybrid approach of using a small amount of looping is faster than the purely vectorized operation.

  • For example:

    N = 10_000
    # Purely vectorized method: Multiply 20XNX15 array against a 1X15 array
    a = np.random.rand(20, N, 15)
    b = np.random.rand(15)
    %timeit np.multiply(b, a)
    
    # Multiply NX15 array against a 1X15 array 20 times
    a = np.random.rand(N, 15)
    b = np.random.rand(15)
    
    %timeit for i in range(20): np.multiply(b, a)
    
    >>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>>
    7.93 ms ± 57.2 µs per loop (mean ± std. dev. of 7 runs, 100 loops each)
    5.43 ms ± 223 µs per loop (mean ± std. dev. of 7 runs, 100 loops each)
    
  • The trend persists until I decrease the array size to below 20X3000X15, at which point the vectorized approach is faster

  • Does anyone know the reason for this? Does numpy have a diminishing return on vectorized operations on large arrays?

  • Is there a better way to do this?

0 Answers
Related