I am currently having difficulties understanding why the following code gets slower after including numba parallelization.
Here is the base code without specifying parallelization:
@njit('f8[:,::1](f8[:,::1], f8[:,::1], f8[:,::1])', fastmath=True)
def fun(A, B, C):
n = A.shape[1]
b00 = B[0,0]
b02 = B[0,2]
out = np.empty((n, 12))
for i in range(n):
ui = A[0,i]
c1 = C[0,i]
c2 = C[1,i]
c3 = C[2,i]
c4 = C[3,i]
out[i, 0] = c1 * b00
out[i, 1] = 0.
out[i, 2] = c1 * (b02-ui)
out[i, 3] = c2 * b00
out[i, 4] = 0.
out[i, 5] = c2 * (b02-ui)
out[i, 6] = c3 * b00
out[i, 7] = 0.
out[i, 8] = c3 * (b02-ui)
out[i, 9] = c4 * b00
out[i, 10] = 0.
out[i, 11] = c4 * (b02-ui)
return out
and here is the parallelized version:
@njit('f8[:,::1](f8[:,::1], f8[:,::1], f8[:,::1])', fastmath=True, parallel=True)
def fun_parallel(A, B, C):
n = A.shape[1]
b00 = B[0,0]
b02 = B[0,2]
out = np.empty((n, 12))
for i in prange(n):
ui = A[0,i]
c1 = C[0,i]
c2 = C[1,i]
c3 = C[2,i]
c4 = C[3,i]
out[i, 0] = c1 * b00
out[i, 1] = 0.
out[i, 2] = c1 * (b02-ui)
out[i, 3] = c2 * b00
out[i, 4] = 0.
out[i, 5] = c2 * (b02-ui)
out[i, 6] = c3 * b00
out[i, 7] = 0.
out[i, 8] = c3 * (b02-ui)
out[i, 9] = c4 * b00
out[i, 10] = 0.
out[i, 11] = c4 * (b02-ui)
return out
Measuring execution times with perfplot and the following code:
B = np.random.rand(3,3)
perfplot.show(
setup=lambda n: (np.random.rand(2, n), np.random.rand(4, n)), # or setup=np.random.rand
kernels=[
lambda A, C: fun(A, B, C),
lambda A, C: fun_parallel(A, B, C),
],
labels=["fun", "fun_parallel"],
n_range=[2**k for k in range(15)],
xlabel="n",
show_progress=False,
)
gives the following performances for a varying size of the arrays.
showing a noticeable increase in time execution with the parallelized version.
Any help on understanding why does this happen is much appreciated.
