The problem:
I have a large workflow which creates at some point an arbitrary number of files per {sample}, named e.g. test1.txt, test2.txt, etc.
I then need to use these files for further processing. The input files for the next rule are then {sample}/test1.txt, {sample}/test2.txt, etc. Thus test1, test2, etc become wildcards.
The data structure is:
---data
---sample1
---test1.txt
---test2.txt
---test3.txt
---sample2
---test1.txt
---test2.txt
Snakefile
I am struggling how snakemake can be used for such problems. I have looked into the function glob_wildcards, but couldn't figure out how to use it.
Intuitively, I would have done something like this:
samples = ['sample1', 'sample2']
rule append_hello:
input:
glob_wildcards('data/{sample}/{id}.txt')
output:
'data/{sample}/{id}_2.txt'
shell:
" echo {input} 'hello' >> {output} "
I have two questions:
- How can this problem be handled in Snakemkae?
- How would you construct a
rule allin order to run this.
Any inputs or any hints into further reading would be appreciated.
Edit
I think it has to do with wildcard constraints. When I run:
assemblies = []
for filename in glob_wildcards(os.path.join("data/{sample}", "{i}.txt")):
assemblies.append(filename)
print(assemblies)
I get two lists where the corresponding index matches:
[['sample1', 'sample1', 'sample1', 'sample2', 'sample2'], ['test1', 'test2', 'test3', 'test5', 'test4']]
Now I basically only need to tell snakemake to use corresponding wildcards values.