I have a dataset as follows:
id email Date_of_purchase time_of_purchase
1 abc@gmail.com 11/10/18 12:10 PM
2 abc@gmail.com 11/10/18 02:11 PM
3 abc@gmail.com 11/10/18 03:14 PM
4 abc@gmail.com 11/11/18 06:16 AM
5 abc@gmail.com 11/11/18 09:10 AM
6 def@gmail.com 11/10/18 12:17 PM
7 def@gmail.com 11/10/18 03:24 PM
8 def@gmail.com 11/10/18 08:16 PM
9 def@gmail.com 11/10/18 09:13 PM
10 def@gmail.com 11/11/18 12:01 AM
I want to calculate the number of transactions made by each email ids within 4 hours. For example, email ids: abc@gmail.com made 3 transactions starting from 11/10/18 12.10 PM to 11/10/18 4.10 PM and made 2 transactions starting from 11/11/18 6.16 AM to 11/11/18 10.16 AM. email ids: def@gmail.com made 2 transactions starting from 11/10/18 12.17 PM to 11/10/18 4.17 PM and made 3 transactions starting from 11/10/18 8.16 PM to 11/11/18 12.16 AM.
My desired output is:
email hour_interval purchase_in_4_hours
abc@gmail.com [11/10/18 12.10 PM to 11/10/18 4.10 PM] 3
abc@gmail.com [11/11/18 6.16 AM to 11/11/18 10.16 AM] 2
def@gmail.com [11/10/18 12.17 PM to 11/10/18 4.17 PM] 2
def@gmail.com [11/10/18 8.16 PM to 11/11/18 12.16 AM] 3
My dataset is having 1000k rows. I am very new in spark. Any help will be highly appreciated. P.S. The time interval can change from 4 hours to 1 hour, 6, hour, 1 day, etc.
TIA.