I have a PySpark dataframe which looks like this, I have a map datatype column Map<Str,Int>
Date Item (Map<Str,int>) Total Items ColA
2021-02-01 Item_A -> 3, Item_B -> 10, Item_C -> 2 15 10
2021-02-02 Item_A -> 1, Item_D -> 5, Item_E -> 7 13 20
2021-02-03 Item_A -> 8, Item_E -> 3, Item_C -> 1 12 30
I want to sum of all the columns including the map column. For map column the sum should be calculated based on keys.
I want something like this:
[[Item_A -> 12, Item_B -> 10, Item_C -> 3, Item_D -> 5, Item_E -> 10], 40, 60]
Not necessarily a list of lists, but I want the sum of the columns.
My approach:
df.rdd.map(lambda x: (1,x[1])).reduceByKey(lambda x,y: x + y).collect()[0][1]