Set global snakemake variables from within a rule

Viewed 125

In snakemake, I can define global variables before declaring my rules. This is nifty when I have a list of files that I need to download via some tool - then I can ensure the name of those files are wildcards for my subsequent steps. For example, when I have a text file (srarunsacc.txt) containing lines of SRA111111,SRA111112,SRA111113 I can do the following (pseudocode):

def read_textfile(name):
....
SRAFILES=read_textfile("srarunsacc.txt")

rule download_srafiles:
    output:
        "%s{srafiles}_1.fastq" % "raw_samples/",
        "%s{srafiles}_2.fastq" % "raw_samples/"
    shell:
        "fasterq-dump {wildcards.srafiles} -O raw_samples"

This ensures that I can then process my downloaded files within snakemake, by linking input and output within rules. But suppose I do not start with the text file srarunsacc.txt and that it is instead generated within the snakemake process. Now I can't define my SRAFILES global variable up top, and I'm stumped at what to do. Ideally, I would be able to make a SRAFILES global variable from within one of my rules, but I haven't managed to do that.


I've created a dummy process meant to mimic my workflow :

rule create_textfile:
    output:
        "list_of_files.txt"
    run:
        files=[f"TESTFILE_{i}.test" for i in range(1,10)]
        files="\n".join(files)
        with open(output[0],"w") as out:
            out.write(files)

In rule create_textfile I create a .txt file that looks like this:

TESTFILE_1.test
TESTFILE_2.test
etc.

Normally, I won't know the exact name of those files, just their extension - which in this case is .test. The next step is :

rule create_files:
    input:
        "list_of_files.txt"
    output:
        "placeholder.plc"
    shell:
        """
        mkdir raw_samples
        cat {input} | parallel touch raw_samples/{{}} >> {output}
        """

In rule create_files I create an empty file from every line of my created .txt file, so the result is this:

raw_samples
    |____TESTFILE_1.test
    |____TESTFILE_2.test
     .etc

Now, in my "experiment", I won't know how the files are named, only their resulting extension, say .test. I want to use those created files as an input to my next rule. I've been using the log output trick to link my rules together. Now I want to process every created file separately. For example:

rule process_files:
input:
    plc="placeholder1.plc",
    sample="raw_samples/{sample}.test"
output:
    "placeholder2.plc"
shell:
    """
    echo {input.sample} >> {output}
    """

The above doesn't work, and gives the error:

Wildcards in input files cannot be determined from output files:'sample'

Which makes sense, as in nowhere above I define output as of raw_samples/{sample}.test So I've tried modifying the rule create_files like this:

 output:
    plc="placeholder1.plc",
    sample="raw_samples/{sample}.test"

That didn't work, and the error was

Not all output, log and benchmark files of rule create_files contain the same wildcards. This is crucial though, in order to avoid that two or more jobs write to the same file.

Tl;dr - A bunch of files (names unknown prior to process, only their extension) is generated via one snakemake rule. I want to process those files with subsequent snakemake rules, but do not know how to link up input-output / set up wildcards.

1 Answers

You are missing a few pieces to get this working the way you want. Snakemake primarily deals with files, so using the list_of_files.txt file is on the right track.

You will need to use a checkpoint on the rule you generate the list of files to determine the sample names. There are lots of questions on checkpoints here you can use as inspiration.

For your "placeholder" outputs, you should change those to directories so snakemake can detect and clean up errors in intermediate steps.

Related