Spark Dataframe vs Pandas Dataframe memory usage comparison

Viewed 367

I'm new to spark, and had written some sample code, to check whether using spark is feasible (reducing memory usage), so I had created a sample dataframe, converted to spark DF, and was comparing the memory usage of both. The sample code is:

emp = [(1, "Smith", -1, "2018", "10", "M", 3000),
       (2, "Rose", 1, "2010", "20", "M", 4000),
       (3, "Williams", 1, "2010", "10", "M", 1000),
       (4, "Jones", 2, "2005", "10", "F", 2000),
       (5, "Brown", 2, "2010", "40", "", -1),
       (6, "Brown", 2, "2010", "50", "", -1)]*800000
empColumns = ["emp_id", "name", "superior_emp_id", "year_joined",
              "emp_dept_id", "gender", "salary"]

df = pd.DataFrame(emp, columns=empColumns)
print(sys.getsizeof(df))

empDF = spark.createDataFrame(data=emp, schema=empColumns)
print(sys.getsizeof(empDF))
print(sys.getsizeof(empDF.collect()))
print(sys.getsizeof(empDF.toPandas()))

The values I'm getting is:

1267200144
48
42915440
1267200144

So now, I have 2 questions:

  1. What is the meaning of sys.getsizeof(empDF)and which part is of 48 bytes here.
  2. Why is there such a difference between spark and pandas df?
0 Answers
Related