Lets take a dataframe of one column with random values. I want to get the rank of all these values which is easy by doing:
df.rank()
But if there are duplicated values you will get a duplicated value also for the rank. For example, for a given list of numbers:
[127.0, 131.856, 132.88, 126.249, 128.417, 124.336, 131.856, 130.624, 147.906, 134.412, 130.735, 133.433, nan, 125.59, 130.211, 133.847, 137.431, 130.0, 127.4, 132.226, 138.134]
the output of the rank function will be:
[4.0, 11.5, 14.0, 3.0, 6.0, 1.0, 11.5, 8.0, 20.0, 17.0, 9.0, 15.0, nan, 2.0, 7.0, 16.0, 18.0, 10.0, 5.0, 13.0, 19.0]
As you can see, the position 1 and 6 are the same and there is no 11 or 12 in the full list. How can we get a rank for these numbers even if it's arbitrary which one goes first?