I have a list var aggList : List[String]= List() the list contains the column names on which aggregation has to be applied.
I generate the dataframe as below:
var df = sc.parallelize(Seq[(Int, Int, String, Int, Int, Int)](
(1234, 1234, "PRM", 2, 1, 1),
(1235, 1234, "PRM", 1239, 2, 10),
(1246, 1234, "PRM", 1234, 5, 15),
(1247, 1234, "PRM", 1254, 20, 12),
(1246, 1234, "PRM", 1234, 5, 13),
(1246, 1234, "SEC", 1234, 7, 15),
(1249, 1234, "SEC", 1234, 20, 1),
(1248, 1234, "SEC", 1234, 2, 2))
).toDF("col1", "col2", "col3", "col4", "col5", "col6")
I need to do df.groupby(col1).agg(sum(aggList))
How do I achieve this?