I have encountered a somewhat unintuitive behavior of keys in data.table package. Here goes an example:
library(data.table)
foo <- data.table(a = c(1:4), b = c(2:5), c = c(3:6), d = c(4:7))
setkey(foo, b)
Then, there is one alarming result of key():
key(foo[, .(mean(c + d)), by = .(b)]) # result is "b".
key(foo[, .(mean(c + d)), by = .(a)]) # result is "a". (!!)
Then, there is another example which produces diffirent, more reasonable results.
foo <- data.table(a = c(4:1), b = c(2:5), c = c(3:6), d = c(4:7))
setkey(foo, b)
key(foo[, .(mean(c + d)), by = .(b)]) # result is "b".
key(foo[, .(mean(c + d)), by = .(a)]) # result is NULL
I admit I'm confused. My lead is this key() somehow checks whether the resulting table needed to be sorted by the elements in by and then assumes it was keyed.
Is it a feature? Is it a bug?