Does NextFlow work for file-oriented cases?

Viewed 288

Reposting from here, hoping for clarification.

Thanks in advance.


I'm completely new to NextFlow and I'm puzzled that I can't do this simple thing, nor find documentation about it: I understand that NF is channel-oriented, but can it process file-oriented cases correctly?

I mean, suppose you have the usual case (see this example, rewrittent from another discussion):

  • process A, creates file a
  • process B, creates file b
  • process C, does something with a+b and creates c (eg, joins a and b)

Now, I delete file a, I expect A and C to be re-executed, with C processing the new a and the existing b and recreating c.

That said,

  • If I do it the regular way, ie, having files into the working dir, deleting working files is unacceptably difficult, since I have to rummage logs and dirs with hashed names until I find what I need. I expect to be able to just delete a or b (or, just touch them).
  • I've tried storeDir and file dates are completely ignored, if a doesn't exist, but c does, A is re-executed but C isn't and the old c is kept.
  • I don't think publishDir would work either, since I'd expect it to work like the first case (for the files remain in the working dir, except for mode='move', which can only be used in a final step).

Am I missing something, or is it that NF isn't a good fit for file-oriented cases like the above?

Moreover, is there a way to run a pipeline only up to a given process (eg, to specify 'A')?

1 Answers

Nextflow processes are executed independently of each other and do not share a common (writable) state. The only way they can communicate is via asynchronous FIFO queues, called channels.

The path input qualifier allows the handling of files in the process execution context. Nextflow will stage the file(s) in the process execution directory so they can be accessed by the script. Note that the path qualifier was introduced in version 19.10.0 as a drop-in replacement for the file qualifier and should be preferred when using a recent version of Nextflow.

In the code you've linked, the storeDir directive is used by each of your three processes. This is far from the 'the usual case' scenario. The directory specified by the storeDir directive is intended as a permanent cache for process results. Generally, these types of processes would have a large one-time cost that you would like to avoid paying again in future runs.

I refactored your example to show that when file 'a' is deleted from the working directory, processes A and C are indeed re-executed:

process A {

    output:
    path 'a.txt' into a_ch

    """
    seq 1 3 > a.txt
    """
}

process B {

    output:
    path 'b.txt' into b_ch

    """
    seq 4 5 > b.txt
    """
}

process C { 

    input:
    path a_txt from a_ch
    path b_txt from b_ch

    output:
    path 'c.txt'
   
    """
    cat "${a_txt}" "${b_txt}" > c.txt
    """
}
$ nextflow run test.nf 
N E X T F L O W  ~  version 20.10.0
Launching `test.nf` [agitated_payne] - revision: baea5be781
executor >  local (3)
[29/4bba82] process > A [100%] 1 of 1 ✔
[d8/978a8d] process > B [100%] 1 of 1 ✔
[14/999791] process > C [100%] 1 of 1 ✔

$ find . -type f -name '*.txt'
./work/14/9997911fcc4587f565822e4c8a238c/c.txt
./work/29/4bba8269b9337f28d20225d605b7cf/a.txt
./work/d8/978a8df7e7e1e5885007ebb0e2915d/b.txt

$ rm ./work/29/4bba8269b9337f28d20225d605b7cf/a.txt

$ nextflow run test.nf -resume
N E X T F L O W  ~  version 20.10.0
Launching `test.nf` [grave_albattani] - revision: baea5be781
executor >  local (2)
[7a/ec8636] process > A [100%] 1 of 1 ✔
[d8/978a8d] process > B [100%] 1 of 1, cached: 1 ✔
[2d/c2d34c] process > C [100%] 1 of 1 ✔

All that said and done, it's still not clear to my why you need to delete files from the working directory in the first place. Without knowing what you're trying to achieve, I'll say avoid touching files in the working directory and instead use the publishDir directive to publish your process results, generally using the 'copy' mode i.e.: publishDir './results', mode: 'copy'. It won't suit all workflows, but it's a good default IMO. Then, if your want to for whatever reason, delete your files from the 'results' folder. When the workflow is resumed, any missing files will be copied across by Nextflow. This strategy will of course keep two copies of your published files on disk. But once you've finished with the workflow, you can simply turf the working directory and just keep the files in 'results' folder.

Related