I will of course have a look at Python multiprocessing with "pool" to solve that, but I'm wondering why is NumPy acting like that.
At best I can see:
- 1 core at 100%
- 2 cores around 70%
- 3 cores around 5%
With similar programs in Julia, there is much more parallelism happening.
Basically, my program is a collection of NumPy vector operations on very large arrays of image and 3D data. The only functions relying on other libraries are Ransac and reports exports with MatPlotLib.
ransac_model, inliers = ransac(data, EllipseModel, min_samples, residual_threshold, max_trials=max_trials)
plt.imshow(self.image)
fig = plt.gcf()
ax = fig.gca()
for x,y,c in zip(annotated_centers_x,annotated_centers_y,annotated_diameters):
circle = plt.Circle((x,y),c,alpha = 0.2,color = 'blue')
ax.add_patch(circle)
plt.savefig(path)
I'm not even sure that the other cores are used by NumpPy. Maybe it is actually these libraries.
I know that it is not the best strategy to run some big programs with NumPy. It would have been wiser to use Julia or TensorFlow/PyTorch tensors. But I have no choice here.
Could you explain me how NumPy behave in terms of parallelism?
I would also be interested if you know how to spot the parts of the code that takes too much time in Python. Something like the profiling library of Julia.
Thanks
