Grateful for some help here. Using Pyspark (cannot use SQL please). So I have a list of tuples stored as RDD Pairs:
[(('City1', '2020-03-27', 'X1'), 44),
(('City1', '2020-03-28', 'X1'), 44),
(('City3', '2020-03-28', 'X3'), 15),
(('City4', '2020-03-27', 'X4'), 5),
(('City4', '2020-03-26', 'X4'), 4),
(('City2', '2020-03-26', 'X2'), 14),
(('City2', '2020-03-25', 'X2'), 4),
(('City4', '2020-03-25', 'X4'), 1),
(('City1', '2020-03-29', 'X1'), 1),
(('City5', '2020-03-25', 'X5'), 15)]
With for example ('City5', '2020-03-25', 'X5') as the Key, and 15 as the value of the last pair.
I would like to obtain the following outcome:
City1, X1, 2020-03-27, 44
City1, X1, 2020-03-28, 44
City5, X3, 2020-03-25, 15
City3, X3, 2020-03-28, 15
City2, X2, 2020-03-26, 14
City4, X4, 2020-03-27, 5
Please notice that the outcome displays:
The Key(s) with the max value for each city (That's the hardest part, to display same city twice if they have similar max(values) in different dates, I'm assuming cannot use ReduceByKey() as Key is not unique, maybe GroupBy() or Filter() ?
In the following sequencing of order/sorting:
- Descending largest value
- Ascending date
- Descending city name (ex: City1)
So I have tried the following code:
res = rdd2.map(lambda x: ((x[0][0],x[0][2]), (x[0][1], x[1])))
rdd3 = res.reduceByKey(lambda x1, x2: max(x1, x2, key=lambda x: x[1]))
rdd4 = rdd3.sortBy(lambda a: a[1][1], ascending=False)
rdd5 = rdd4.sortBy(lambda a: a[1][0])
Although it does give me the cities with the max value, it doesn't return the same city twice (because reduced by Key: City) if 2 cities has similar max value in 2 different dates.
I hope its clear enough, any precision please ask! Thanks so much!