I'm new to spark, and had written some sample code, to check whether using spark is feasible (reducing memory usage), so I had created a sample dataframe, converted to spark DF, and was comparing the memory usage of both. The sample code is:
emp = [(1, "Smith", -1, "2018", "10", "M", 3000),
(2, "Rose", 1, "2010", "20", "M", 4000),
(3, "Williams", 1, "2010", "10", "M", 1000),
(4, "Jones", 2, "2005", "10", "F", 2000),
(5, "Brown", 2, "2010", "40", "", -1),
(6, "Brown", 2, "2010", "50", "", -1)]*800000
empColumns = ["emp_id", "name", "superior_emp_id", "year_joined",
"emp_dept_id", "gender", "salary"]
df = pd.DataFrame(emp, columns=empColumns)
print(sys.getsizeof(df))
empDF = spark.createDataFrame(data=emp, schema=empColumns)
print(sys.getsizeof(empDF))
print(sys.getsizeof(empDF.collect()))
print(sys.getsizeof(empDF.toPandas()))
The values I'm getting is:
1267200144
48
42915440
1267200144
So now, I have 2 questions:
- What is the meaning of
sys.getsizeof(empDF)and which part is of 48 bytes here. - Why is there such a difference between spark and pandas df?