I used to make some parallelizable computations using python on a server, now I want to migrate to an HPC working on SLURM, but I believe I have some conceptual issues.
My computation needs matrices and parameters (as dictionary) that are same for all runs, and some input parameters that differ. So it basically looks like:
def evaluate(inp, matrices, params):
.....
return output
load matrices, create params dict...
inputs = [(k, matrices, params) for k in input_parameters]
with multiprocessing.Pool(processes=n) as pool:
results =pool.starmap(evaluate,inputs)
save_results(results)
and a bash script:
#!/bin/bash
#SBATCH -J ****
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=40
#SBATCH --mem=500M
#SBATCH --time=1:00:00
#SBATCH --partition=standard
module load python
source $HOME/env/bin/activate
python3 main.py
So my problem here is, on a single node, even though processes are active, they get %0 CPU, and waits until SLURM exits the job due to over-time, and when I have tried to use multiprocessing.get_context.Pool instead of multiprocessing.Pool, the code runs, and never finishes. I don't quite see what the problem can be?
Also if I want to execute this program on multiple nodes, I understand that I need to bind MPI's, and mpi4py or ray can be used for this purpose. Is this also the case for a single node job?