I have a hive table like the following. I want to assign the row number for user consecutive data. When the data is not consecutive, the row number will reset to 1.
| Date | User |
|------------|------|
| 2022-05-01 | A |
| 2022-05-02 | A |
| 2022-05-03 | A |
| 2022-05-05 | A |
| 2022-05-06 | A |
| 2022-05-01 | B |
| 2022-05-03 | B |
| 2022-05-04 | B |
| 2022-05-05 | B |
| 2022-05-06 | B |
| 2022-05-01 | C |
| 2022-05-02 | C |
| 2022-05-04 | C |
| 2022-05-05 | C |
| 2022-05-06 | C |
Here is the output I want.
| Date | User | Rank |
|------------|------|------|
| 2022-05-01 | A | 1 |
| 2022-05-02 | A | 2 |
| 2022-05-03 | A | 3 |
| 2022-05-05 | A | 1 |
| 2022-05-06 | A | 2 |
| 2022-05-01 | B | 1 |
| 2022-05-03 | B | 1 |
| 2022-05-04 | B | 2 |
| 2022-05-05 | B | 3 |
| 2022-05-06 | B | 4 |
| 2022-05-01 | C | 1 |
| 2022-05-02 | C | 2 |
| 2022-05-04 | C | 1 |
| 2022-05-05 | C | 2 |
| 2022-05-06 | C | 3 |
I couldn't figure out with row_number, lag and other udfs.