How to make numba @jit use all cpu cores (parallelize numba @jit)

Viewed 12380

I am using numbas @jit decorator for adding two numpy arrays in python. The performance is so high if I use @jit compared with python.

However it is not utilizing all CPU cores even if I pass in @numba.jit(nopython = True, parallel = True, nogil = True).

Is there any way to to make use of all CPU cores with numba @jit.

Here is my code:

import time                                                
import numpy as np                                         
import numba                                               

SIZE = 2147483648 * 6                                      

a = np.full(SIZE, 1, dtype = np.int32)                     

b = np.full(SIZE, 1, dtype = np.int32)                     

c = np.ndarray(SIZE, dtype = np.int32)                     

@numba.jit(nopython = True, parallel = True, nogil = True) 
def add(a, b, c):                                          
    for i in range(SIZE):                                  
        c[i] = a[i] + b[i]                                 

start = time.time()                                        
add(a, b, c)                                               
end = time.time()                                          

print(end - start)                                        
2 Answers

For the sake of completeness, in year 2018 (numba v 0.39) you can just do

from numba import prange

and replace range with prange in your original function definition, that's it.

That immediately makes CPU utilization 100% and in my case speeds things up from 2.9 to 1.7 seconds of runtime (for SIZE = 2147483648 * 1, on machine with 16 cores 32 threads).

More complex kernels one often can speed up even more by passing in fastmath=True.

Related