I need help in using windowing function in pyspark. I have a following type of dataframe
| id | EventDate |
|---|---|
| 1 | 01/01/2022 |
| 1 | 02/01/2022 |
| 1 | 03/01/2022 |
| 1 | 04/01/2022 |
| 1 | 05/01/2022 |
| 1 | 06/01/2022 |
| 1 | 08/01/2022 |
| 1 | 09/01/2022 |
| 1 | 10/01/2022 |
| 1 | 11/01/2022 |
| 1 | 12/01/2022 |
| 1 | 13/01/2022 |
| 1 | 14/01/2022 |
| 2 | 01/01/2022 |
| 2 | 02/01/2022 |
| 2 | 03/01/2022 |
| 2 | 04/01/2022 |
| 2 | 05/01/2022 |
| 2 | 06/01/2022 |
| 2 | 07/01/2022 |
| 2 | 08/01/2022 |
| 2 | 09/01/2022 |
| 2 | 10/01/2022 |
| 2 | 11/01/2022 |
| 2 | 12/01/2022 |
| 2 | 13/01/2022 |
| 2 | 14/01/2022 |
My goal is for each id to window the data in 7 days range and assign unique value for each range. As, for example,in the following table:
| id | EventDate | Week number |
|---|---|---|
| 1 | 01/01/2022 | 1 |
| 1 | 02/01/2022 | 1 |
| 1 | 03/01/2022 | 1 |
| 1 | 04/01/2022 | 1 |
| 1 | 05/01/2022 | 1 |
| 1 | 06/01/2022 | 1 |
| 1 | 08/01/2022 | 2 |
| 1 | 09/01/2022 | 2 |
| 1 | 10/01/2022 | 2 |
| 1 | 11/01/2022 | 2 |
| 1 | 12/01/2022 | 2 |
| 1 | 13/01/2022 | 2 |
| 1 | 14/01/2022 | 2 |
| 2 | 01/01/2022 | 1 |
| 2 | 02/01/2022 | 1 |
| 2 | 03/01/2022 | 1 |
| 2 | 04/01/2022 | 1 |
| 2 | 05/01/2022 | 1 |
| 2 | 06/01/2022 | 1 |
| 2 | 07/01/2022 | 1 |
| 2 | 08/01/2022 | 2 |
| 2 | 09/01/2022 | 2 |
| 2 | 10/01/2022 | 2 |
| 2 | 11/01/2022 | 2 |
| 2 | 12/01/2022 | 2 |
| 2 | 13/01/2022 | 2 |
| 2 | 14/01/2022 | 2 |
Please Note that it is not fact that each id has all EventDatein row, i.e. the id = 1 has no row for EventDate=07/01/2022, nevertheless the function should understand that the second week starts from 08/01/2022
Is there any solution for this problem?