I got a pyspark dataframe that looks like:
| id | score |
|---|---|
| 1 | 0.5 |
| 1 | 2.5 |
| 2 | 4.45 |
| 3 | 8.5 |
| 3 | 3.25 |
| 3 | 5.55 |
And I want to create a new column rank based on the value of the score column in incrementing order meaning the highest value will have the rank 0 and restarting the count by the id column.
Something like this:
| id | value | rank |
|---|---|---|
| 1 | 2.5 | 0 |
| 1 | 0.5 | 1 |
| 2 | 4.45 | 0 |
| 3 | 8.5 | 0 |
| 3 | 5.55 | 1 |
| 3 | 3.25 | 2 |
Thanks in advance!