I am building a data processing pipeline. The data is quite large: a data frame representing sensor data sampled at a high frequency. During the pipeline, I have an intermediate result which is a transformation of the data which is needed in subsequent transformations. Using Dask, I found that the intermediate transformation has to be re-computed in each subsequent transformation.
Is there a way to persist the intermediate result on disk? I am aware of .persist(), but this keeps the result on memory which is not an option due to the large size of my data.