I have pipeline that takes in input different species. If the value of the {species} wildcard is "mouse" or "human", I need to do some pre-processing common to both species and execute some rules, otherwise execute another set of rules. This is pseudocode of what I'm trying to achieve:
SPECIES = ['mouse', 'human', 'pig']
rule all:
input:
expand('{species}.done', species=SPECIES),
if wildcards.species in ['mouse', 'human']:
rule prepare_data:
output:
'some.data'
rule mouse_human:
input:
'some.data',
output:
'{species}.tmp',
else:
rule animal:
# Note file "some.data" is not needed
output:
'{species}.tmp',
rule done:
input:
'{species}.tmp',
output:
'{species}.done',
That is: If {species} is "mouse" or "human", run rule prepare_data (only once) and then run rule mouse_human twice, once for human once for mouse. If {species} is "pig" or something else run only rule animal.
The pseudocode above won't run because if wildcards.species in ['mouse', 'human']: is not valid syntax. How can I do that?
A possible solution would be this:
rule prepare_data:
output:
'some.data',
rule species:
input:
'some.data',
output:
'{species}.tmp',
run:
if wildcards.species in ['mouse', 'human']:
` # Do human/mouse stuff using "some.data" and output {species}.txt
else:
# Do other stuff and output {species}.tmp, ignore "some.data"
However, rule prepare_data would always run even if the user's input data does not include "mouse" or "human". This is wasteful and I would like to avoid it.