When I cache a skewed table in spark, it takes 4x times as compared to caching a uniformly distributed dataset. What I do is repartition the dataset uniformly before caching. I want to know why this happens, isn't the dataset already in memory, how does caching take so much time? I verified that even though the dataset was skewed, caching never overflowed to disk and was always accommodated in memory.
Sample Example
Dataset skewedDs = ds.read();
ds.cache();
ds.first();
Above takes 2.5mins to evaluate
Dataset skewedDs = ds.read();
uniformDS = ds.repartition(n)
uniformDS.cache();
uniformDS.first();
Above takes 40secs to evaluate. As you can see, simply caching a skewed dataset takes a huge chunk of time.