Performance comparison between Pivot vs Filters+Joins

Viewed 96

I would like to ask you opinion about the following question:

From a computational effort point of view, in a pyspark environment, would you advise to use a pivot function or a series of filtered tables and joins to carry out the same result? And why?

From my point of view, despite coding a pivot should be faster, I would prefer a series of joins on filtered results because it wouldn’t require any transposition on the dataframe, but I am here to ask you opinion in order to learn more about it!

Thanks

0 Answers
Related