What is the best data structure for scipy.cKDTree.query?

Viewed 113

I am using scipy.cKDTree.query(data, n_jobs=4) in conjunction with export OMP_NUM_THREADS=4, on a machine with 4 threads per core. I would have expected there to be 400% CPU usage when the script is running, however, only around 200% seems to be used with numpy arrays. With the zipped data 400% is reached but only intermittently. I wondered what the best data structure is for scipy.cKDTree.query?

This is the simple shell script (script.sh):

#!/bin/sh
export OMP_NUM_THREADS=4
python -u stackOverflowExample.py > log.python 2>&1 &

This is the simple python code (stackOverflowExample.py) for the zipped data (reached 400% CPU but only briefly three times and returned to 100% CPU in between):

from scipy import spatial
import numpy as np

x, y, z = np.mgrid[2.111:400, 2.222:640, 2.543:48]
data = list(zip(x.ravel(), y.ravel(), z.ravel()))

x, y, z = np.mgrid[1.1:400, 1.5:640, 1.125:48]
data2 = list(zip(x.ravel(), y.ravel(), z.ravel()))

tree = spatial.cKDTree(data)

for i in range(3):

    distance, index = tree.query(data2,n_jobs=4)

I also tried a similar code with numpy arrays (reaches around 200% CPU and largely constant):

from scipy import spatial
import numpy as np
import time

x, y, z = np.mgrid[2.111:400, 2.222:640, 2.543:48]
data3 = np.array([x.ravel(), y.ravel(), z.ravel()])

x, y, z = np.mgrid[1.1:400, 1.5:640, 1.125:48]
data4 = np.array([x.ravel(), y.ravel(), z.ravel()])
data5=data4[:,:data3.shape[1]]

tree = spatial.cKDTree(data3)

for i in range(50):

    distance, index = tree.query(data5,n_jobs=4)

  • Python version: 3.6.5
  • Scipy version: 1.1.0
  • Numpy version: 1.14.3
0 Answers
Related