I am using scipy.cKDTree.query(data, n_jobs=4) in conjunction with export OMP_NUM_THREADS=4, on a machine with 4 threads per core. I would have expected there to be 400% CPU usage when the script is running, however, only around 200% seems to be used with numpy arrays. With the zipped data 400% is reached but only intermittently. I wondered what the best data structure is for scipy.cKDTree.query?
This is the simple shell script (script.sh):
#!/bin/sh
export OMP_NUM_THREADS=4
python -u stackOverflowExample.py > log.python 2>&1 &
This is the simple python code (stackOverflowExample.py) for the zipped data (reached 400% CPU but only briefly three times and returned to 100% CPU in between):
from scipy import spatial
import numpy as np
x, y, z = np.mgrid[2.111:400, 2.222:640, 2.543:48]
data = list(zip(x.ravel(), y.ravel(), z.ravel()))
x, y, z = np.mgrid[1.1:400, 1.5:640, 1.125:48]
data2 = list(zip(x.ravel(), y.ravel(), z.ravel()))
tree = spatial.cKDTree(data)
for i in range(3):
distance, index = tree.query(data2,n_jobs=4)
I also tried a similar code with numpy arrays (reaches around 200% CPU and largely constant):
from scipy import spatial
import numpy as np
import time
x, y, z = np.mgrid[2.111:400, 2.222:640, 2.543:48]
data3 = np.array([x.ravel(), y.ravel(), z.ravel()])
x, y, z = np.mgrid[1.1:400, 1.5:640, 1.125:48]
data4 = np.array([x.ravel(), y.ravel(), z.ravel()])
data5=data4[:,:data3.shape[1]]
tree = spatial.cKDTree(data3)
for i in range(50):
distance, index = tree.query(data5,n_jobs=4)
- Python version: 3.6.5
- Scipy version: 1.1.0
- Numpy version: 1.14.3