Assuming that we have one node with 12 cores. What are the differences between:
- Run one MPI process to manage 12 threads for each core.
- Just run 12 MPI processes for each core.
The former communicate via shared memory, and the latter communicate via IPC. So, which one is faster? are the differences negligible or significant?