For python dataframe, info() function provides memory usage. Is there any equivalent in pyspark ? Thanks
For python dataframe, info() function provides memory usage. Is there any equivalent in pyspark ? Thanks
I have something in mind, its just a rough estimation. as far as i know spark doesn't have a straight forward way to get dataframe memory usage, But Pandas dataframe does. so what you can do is.
sample = df.sample(fraction = 0.01)pdf = sample.toPandas()pdf.info()As per the documentation:
The best way to size the amount of memory consumption a dataset will require is to create an RDD, put it into cache, and look at the “Storage” page in the web UI. The page will tell you how much memory the RDD is occupying.
To estimate the memory consumption of a particular object, use SizeEstimator’s estimate method. This is useful for experimenting with different data layouts to trim memory usage, as well as determining the amount of space a broadcast variable will occupy on each executor heap.
You can persist dataframe in memory and take action as df.count(). You would be able to check the size under storage tab on spark web ui.. let me know if it works for you.
How about below? It's in KB, X100 to get the estimated real size.
df.sample(fraction = 0.01).cache().count()