CPU/Threads usage on M1 Pro (Apple Silicon) using openMP

Viewed 429

hope someone knows the answer to this...

I have a code that compiles perfectly well with openMP (it uses libsharp). However, I am finding it impossible to make the M1 Pro chip use all the 8 or 10 cores I have.

I am setting the threads variable correctly as export OMP_NUM_THREADS=10 such that the code correctly identifies it's supposed to be running with 10 threads (see image below showing a print-screen from my activity monitor):

Activity Monitor Print Screen

Print screen is showing that the code is compiled for Apple Silicon, uses 10 threads but not much of the CPU available.

Does anyone know how to properly compile/set the number of threads such that all the cores will be used?

This is trivial in x86 architectures.

1 Answers

Not really an answer, but long for a comment...

If both LLVM and GCC behave the same then it's not an OpenMP runtime issue. (And your monitor output shows that the correct number of threads have been created). I'm also not certain that it's really an Arm issue. Are you comparing with an Apple x86 machine (so running the same operating system), or with a Linux x86 system? The scheduling decisions of the two OSes are likely different, and (for instance) MacOS has no interface to bind threads to logicalCPUs. As well as that, there's the issue of having some fast and some slow cores. That could mean that statically scheduled loops are inefficient. I'm also confused by the fact that you arm to show multiple instances of your code running at the same time, so you are explicitly causing over-subscription of the logicalCPUs...

Related