for the project I am working on I use pyspark on Databricks, I need to transform data from different sources and tables. Because of bad data quality, there are a lot of joins to do to get the data from different sources like a puzzle. As the code increases in size, calculations are getting longer and longer. Even sometimes calculations are stuck for several tens of minutes without any changes.
I found that by saving regularly the dataframe in parquet format in dbfs, and then read this save from parquet just after having saved it before continuing calculation hugely increases the computation speed.
So I would like to continue to do it this way as I don't see any drawback so far.
However as this should be automated and used in production, I wanted to understand if there are drawbacks of doing so I should be aware of ? Saving the dataframe on dbfs is only used to speedup computation, at the end this save is of no use and can be deleted.
Thanks