Snakemake: Cluster multiple jobs together

Viewed 781

I have a pretty simple snakemake pipeline that takes an input file does three subsequent steps to produce one output. Each individual job is very quick. Now I want to apply this pipeline to >10k files on an SGE cluster. Even if I use group to have one job for each three rules per input file, I would still submit >10k cluster jobs. Is there a way to instead submit limited number of cluster jobs (lets say 100) and distribute all tasks equally between them?

An example would be something like

rule A:
        input: {prefix}.start
        output: {prefix}.A
        group "mygroup"

rule B:
        input: {prefix}.A
        output: {prefix}.B
        group "mygroup"

rule C:
        input: {prefix}.B
        output: {prefix}.C
        group "mygroup"

rule runAll:
        input: expand("{prefix}.C", prefix = VERY_MANY_PREFIXES)

and then run it with snakemake --cluster "qsub <some parameters>" runAll

2 Answers

You could process all the 10k files in the same rule using a for loop (not sure if this is what Manavalan Gajapathy has in mind). For example:

rule A:
    input:
        txt= expand('{prefix}.start', prefix= PREFIXES),
    output:
        out= expand('{prefix}.A', prefix= PREFIXES),
    run:
        io= zip(input.txt, output.out)
        for x in io:
            shell('some_command %s %s' %(x[0], x[1]))

and the same for rule B and C.

Look also at snakemake local-rules

The only solution I can think of would be to declare rules A, B, and C to be local rules, so that they run in the main snakemake job instead of being submitted as a job. Then you can break up your runAll into batches:

rule runAll1:
        input: expand("{prefix}.C", prefix = VERY_MANY_PREFIXES[:1000])

rule runAll2:
        input: expand("{prefix}.C", prefix = VERY_MANY_PREFIXES[1000:2000])

rule runAll3:
        input: expand("{prefix}.C", prefix = VERY_MANY_PREFIXES[2000:3000])

...etc

Then you submit a snakemake job for runAll1, another for runAll2, and so on. You do this fairly easily with a bash loop:

for i in {1..10}; do sbatch [sbatch params] snakemake runAll$i; done;

Another option which would be more scalable than creating multiple runAll rules would be to have a helper python script that does something like this:

import subprocess

for i in range(0, len(VERY_MANY_PREFIXES), 1000):
    subprocess.run(['sbatch', 'snakemake'] + ['{prefix}'.C for prefix in VERY_MANY_PREFIXES[i:i+1000]])
Related