I have the below PySpark dataframe. column_2 is of complex data type array<map<string,bigint>>
Column_1 Column_2 Column_3
A [{Mat=7},{Phy=8}] ABC
A [{Mat=7},{Phy=8}] CDE
B [{Mat=6},{Phy=7}] ZZZ
I have to group by on column 1 and column 2 and get the minimum aggregate of column 3.
The problem is when I try to group by column 1 and column 2 it's giving me an error
cannot be used as grouping expression because the data type is not an orderable data type
Is there a way to include this column in group by or to aggregate it in some way. The values in column_2 will always be same for a key value in column_1
Expected output:
Column_1 Column_2 Column_3
A [{Mat=7},{Phy=8}] ABC
B [{Mat=6},{Phy=7}] ZZZ
Is it possible to do a collect list of all value in aggregate function and flatten it and remove duplicates?