I currently build lags and leads before I run regressions. This feels a bit clumsy (even compared to Stata where lagging happens automatically by category after xtset) and I'd like to lag directly in the regression formula. However, I do not know how to combine the by = part of data.table with its shift() function inside a regression command such as lm().
Example:
Create some data (the NA's just make sure that lagging and lagging by category don't coincidentally create the same result...):
library(data.table)
set.seed(123)
DT <- data.table(id = c(rep("A", 4), rep("B", 3)),
y = rnorm(7),
dummy = c(0, 1, 0, NA, NA, 1, 0))
Creating lags before running regressions works of course but is tedious and (with many lags) clutters up the data:
DT[, mylag := shift(dummy, fill=0), by = id]
shift() works inside lm() but I cannot shift by category so the results differ from those of the previously created lag:
lm(y ~ mylag, data=DT)
lm(y ~ shift(dummy), data=DT)
Which gets me back to my question: How can I call shift() by category inside a regression?