I have a DataFrame like this (but much larger):
id start end
0 10 20
1 11 13
2 14 18
3 22 30
4 25 27
5 28 31
I am trying to efficiently merge overlapping intervals in PySpark, while saving in a new column 'ids', which intervals were merged, so that it looks like this:
start end ids
10 20 [0,1,2]
22 31 [3,4,5]
Visualisation:
from:
Can I do this without using an udf?
edit: the order of id and start are not necessarily the same.

