Consider following snippet:
library("data.table")
dT <- data.table(keyCol=sample(x = c("A", "B", "C"), size = 20, replace = TRUE),
valCol=rpois(n = 20, lambda=10))
# head(dT)
keyCol valCol
<chr> <int>
A 11
C 14
C 9
B 9
C 11
C 10
I want to calculate some aggregates (unique count of valCol, number of rows) grouping keyCol column:
res <- dT[, .(unique__=length(unique(valCol)), count__=.N), by="keyCol"]
# res
keyCol unique__ count__
<chr> <int> <int>
A 4 4
C 5 8
B 6 8
But I want to calculate these aggregates conditionally, i.e. sometimes I want unique only, sometimes I want count only, and sometimes I want both. One possible solution is to use multiple if else conditions:
getCount <- TRUE
getUnique <- TRUE
aggColName <- "valCol"
if(getCount & getUnique){
res <- dT[, .(unique__=length(unique(valCol)), count__=.N), by="keyCol"]
} else if (getCount){
res <- dT[, .(count__=.N), by="keyCol"]
} else {
res <- dT[, .(unique__=length(unique(valCol))), by="keyCol"]
}
# res
keyCol unique__ count__
<chr> <int> <int>
A 4 4
C 5 8
B 6 8
Coming from Python background, above if else ladder seems like writing bunch of redundant codes. For example, on pandas, we can pass tuple containing the column names, and aggregate functions, or we can unpack a dictionary containing column name, and aggregate functions.
Is there simple way of doing the same using data.table, i.e. passing column names, and aggregates function using some variables, or passing them conditionally?