Couldn't really get a straight answer from the net. Consider the following data scenario: I have data which contains user_id and timestamps of user activity:
val bigData = Seq( ( "id3",12),
("id1",55),
("id1",59),
("id1",50),
("id2",51),
("id3",52),
("id2",53),
("id1",54),
("id2", 34)).toDF("user_id", "ts")
So the original DataFrame looks like this:
+-------+---+
|user_id| ts|
+-------+---+
| id3| 12|
| id1| 55|
| id1| 59|
| id1| 50|
| id2| 51|
| id3| 52|
| id2| 53|
| id1| 54|
| id2| 34|
+-------+---+
and this is what I will write to HDFS\S3 for example.
However I can't save the data grouped by user for instance like this:
bigData.groupBy("user_id").agg(collect_list("ts") as "ts")
Which result in:
+-------+----------------+
|user_id| ts|
+-------+----------------+
| id3| [12, 52]|
| id1|[55, 59, 50, 54]|
| id2| [51, 53, 34]|
+-------+----------------+
I can get a decisive answer on which method will get better storage/compression on the filesystem. The grouped approach looks (intuitively) better storage/comperssion wise.
Anyone know if there is absolute approch or know any benchmarks or articles on this subject?