Ranking Function Output

Viewed 32

The book R for Data Science goes over ranking creation functions, and I'm having trouble understanding the examples even after looking at the documentation.

Here is the example:

y <- c(1, 2, 2, NA, 3, 4)
min_rank(y)
#> [1]  1  2  2 NA  4  5
min_rank(desc(y))
#> [1]  5  3  3 NA  2  1 
row_number(y)
#> [1]  1  2  3 NA  4  5 
dense_rank(y)
#> [1]  1  2  2 NA  3  4 
percent_rank(y)
#> [1] 0.00 0.25 0.25   NA 0.75 1.00 
cume_dist(y)
#> [1] 0.2 0.6 0.6  NA 0.8 1.0

Questions: min_rank: - where is the 5 from? and why isn't NA last? min_rankd(desc()) -- why are there 2 3s and not 2 2's? row_number: still confused on NA positoning, and wouldn't there be 6 rows?

1 Answers

5 is there because while the two 2s are tied, the second 2 still counts as an extra rank.

It appears that min_rank and row_number are simply convenience analogs for rank with less customizability as to how NAs are handled. Instead of min_rank, you can use:

rank(y, ties.method = "min", na.last = c(TRUE, FALSE, NA, "keep"))

The rank documentation says the following:

na.last:

for controlling the treatment of NAs. If TRUE, missing values in the data are put last; if FALSE, they are put first; if NA, they are removed; if "keep" they are kept with rank NA.

So it appears that the default for min_rank and row_number is just keep, but you can customize that to TRUE if you like.

As for your second question about desc:

desc outputs a negative version of the numeric vector as such:

[1] -1 -2 -2 NA -3 -4

So -1 is the highest (marked as 5), -2 is tied for second highest (marked as 3), et cetera. Let me know if this answers your question.

Related