Running parallel python programs on HPC

Viewed 131

I used to make some parallelizable computations using python on a server, now I want to migrate to an HPC working on SLURM, but I believe I have some conceptual issues.

My computation needs matrices and parameters (as dictionary) that are same for all runs, and some input parameters that differ. So it basically looks like:

def evaluate(inp, matrices, params):
    .....
    return output

load matrices, create params dict...
inputs = [(k, matrices, params) for k in input_parameters]
with multiprocessing.Pool(processes=n) as pool:
    results =pool.starmap(evaluate,inputs)
save_results(results)

and a bash script:

#!/bin/bash
#SBATCH -J ****     
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=40
#SBATCH --mem=500M              
#SBATCH --time=1:00:00 
#SBATCH --partition=standard
module load python
source $HOME/env/bin/activate
python3 main.py

So my problem here is, on a single node, even though processes are active, they get %0 CPU, and waits until SLURM exits the job due to over-time, and when I have tried to use multiprocessing.get_context.Pool instead of multiprocessing.Pool, the code runs, and never finishes. I don't quite see what the problem can be?

Also if I want to execute this program on multiple nodes, I understand that I need to bind MPI's, and mpi4py or ray can be used for this purpose. Is this also the case for a single node job?

0 Answers
Related