I want to group my dataframe elements, basing on two columns in both directions. This is a sample of used dataframe
val columns = Seq("src","dst")
val data = Seq(("A", "B"), ("B", "C"), ("C", "A"),("A", "B"), ("B", "A"), ("B", "A"),("A", "C"), ("B", "A"), ("C", "D"),("D", "C"), ("A", "C"), ("C", "A"))
val rdd = spark.sparkContext.parallelize(data)
val dff = spark.createDataFrame(rdd).toDF(columns:_*)
When I use simple groupBy on two columns I get this result
dff.groupBy("src","dst").count().show()
+---+---+-----+
|src|dst|count|
+---+---+-----+
| B| C| 1|
| D| C| 1|
| A| C| 2|
| C| A| 2|
| C| D| 1|
| B| A| 3|
| A| B| 2|
+---+---+-----+
I want group columns where src and dst are the same in the other direction(for example grouping A,C and C,A together, A,B and B,A together...).
The desired result is like that
+---+---+-----+
|src|dst|count|
+---+---+-----+
| B| C| 1|
| D| C| 2|
| A| C| 4|
| B| A| 5|
+---+---+-----+
Any solutions ?