I am doing a group by over a column in a pyspark dataframe and doing a collect list on another column to get all the available values for column_1. As below.
Column_1 Column_2
A Name1
A Name2
A Name3
B Name1
B Name2
C Name1
D Name1
D Name1
D Name1
D Name1
The output that i get is a collect list of column_2 with column_1 grouped.
Column_1 Column_2
A [Name1,Name2,Name3]
B [Name1,Name2]
C [Name1]
D [Name1,Name1,Name1,Name1]
Now when all the values within collect list are same, i just want to display it only once and not four times. Below is the expected output.
Expected output:
Column_1 Column_2
A [Name1,Name2,Name3]
B [Name1,Name2]
C [Name1]
D [Name1]
Is there a way to do this in pyspark?