Toy example:
library(data.table)
set.seed(1)
n_people <- 100
groups <- c("A", "B", "C")
example_table <- data.table(person_id=seq_len(n_people),
group_2010=sample(groups, n_people, TRUE),
group_2011=sample(groups, n_people, TRUE))
## Error-prone and requires lots of typing -- programmatic alternative?
transition_probs <- example_table[, list(pr_A_2011=mean(group_2011=="A"),
pr_B_2011=mean(group_2011=="B"),
pr_C_2011=mean(group_2011=="C")),
by=group_2010]
transition_probs # Essentially a transition matrix giving Pr[group_2011 | group_2010]
# group_2010 pr_A_2011 pr_B_2011 pr_C_2011
# 1: A 0.1481481 0.5185185 0.3333333
# 2: B 0.3684211 0.3947368 0.2368421
# 3: C 0.3142857 0.3142857 0.3714286
The "manual" approach above is fine when the groups are A, B, C, but gets messy if there are more groups (or if we just have the groups vector but don't know ahead of time what it contains).
Is there a "data.table way" to compute the transition_probs object in my example code above? Can list(pr_A_2011=...) be replaced with something programmatic?
My concern is that, if I add a group D, I will have to edit the code in multiple places, notably by typing pr_D_2011=mean(group_2011=="D").