Using Snakemake to handle SLURM directives without running processes on login node

Viewed 36

I'm trying to use Snakemake to handle the running of an RNASeq pipeline on a HPC using SLURM to schedule jobs. I read that Snakemake has native support to schedule jobs using SBATCH on a cluster using snakemake --cluster, however this creates a Catch-22:

If you run snakemake --cluster from a login node, the Snakemake process it generates continues running on the login node until the pipeline stops running, which degrades performance of the login node and is against the rules of the HPC cluster.

If you run snakemake --cluster from a compute node, you are dispatching SBATCH commands from within an interactive session and on a compute node, while the rules of the cluster dictate they must be dispatched from the login node only.

Is there any way to use Snakemake to handle parallelization on a SLURM-managed cluster while 1) not running any long-term processes on the login node and 2) dispatching all jobs from the login node? My naive idea is to run Snakemake on a compute node, and then somehow get back to the login node before issuing the SBATCH commands, but I have no idea if this is even possible.

Thanks in advance for your time :)

2 Answers

Running snakemake on the login node is unlikely to cause problems because all this process does is periodically check the job status (assuming that the local rules are not resource-intensive). To be sure one can check with the cluster support team (I did and they said it was fine).

I will echo that the main snakemake process is not compute or memory intensive once it gets going. If you are concerned about swamping resources, running nice snakemake will make sure other processes get priority. Still, I've heard of clusters where basically no commands are allowed on the head node and the policy is enforced. If the sys admins are not willing to make an exception, you should ask them how to submit jobs from worker nodes; you certainly aren't the only person using a workflow manager at your organization. Some clusters allow direct sbatch submission but you may also have to ssh back to the login node.

Related