Functional data processing in clojure

Viewed 151

I'm looking for some architectural solution to the problem I'm facing. I am struggling with a huge pipeline that processes data (which is a very complex structure). Generally, it can be illustrated as:

(defn process [data]
  (-> data
    do-something-1
    do-something-2
    do-something-3
    do-something-4
    ...
    do-something-20
)

where each of the do-something-* functions might be similarly complex. My problem is that there is a lot of coupling between functions in such a processing chain. For example do-something-3add something to data that later is required by do-something-9 which adds something else required by do-something-18 and so on. So essentially data is enriched by all those functions when cascading down the threading macro. It's very hard to keep the track of what is happening and when. Holding the whole processing chain in my head is just too much of a cognitive load (or at least I have too little RAM in my head). How do handle such cases? I get that there is no silver bullet but maybe there is something I'm missing (I started to learn clojure few months ago).

1 Answers

I have never liked long pipelines as you describe because of the difficulties you are encountering.

As a first step, you might just break up the pipeline and give a name (even if temporary) to each stage:

(defn process [data]
  (let [x01 (do-something-1 data)
        x02 (do-something-2 x01)
        x03 (do-something-3 x02)
        x04 (do-something-4 x03)
        ; ...
        x20 (do-something-20 x19)]
    x20))

Then, you could add debug and/or validation statements between as necessary. I would also suggest using Plumatic Schema to document the input/output of each do-something-* function (or maybe Malli; I don't like spec). Hopefully the function names are also descriptive; you could improve those as another plus.

You could also group the functions in a hierarchy instead of a just a linear chain:

A
 - A1
 - A2
 - A3
 - A4
B
 - B1 
 - B2
 - B3
C 
 ...etc...

So the top-level pipeline is only 3-4 calls, each of which may have 2-5 sub-calls.

Of course, I hope you have nice unit tests for each function at all levels to document the expected behavior and typical inputs/outputs.

Related