I have a mixed effect model trained on one dataset, and then used to generate predicted values for another dataset. How do I obtain the marginal R2 for the new predicted values?
As an example, let's break the sleepstudy data from the lme4 package into two subsamples, keeping participants grouped together:
library(lme4)
set.seed(33)
dataA_list <- lapply(split(seq(1:nrow(sleepstudy)), sleepstudy$Subject), function(x) sample(x, floor(.5*length(x))))
dataB_list <- mapply(function(x,y) setdiff(x,y), x = split(seq(1:nrow(sleepstudy)), sleepstudy$Subject), y = dataA_list)
dataA <- sleepstudy[unlist(dataA_list),]
dataB <- sleepstudy[unlist(dataB_list),]
I trained a mixed effects model on the first dataset (dataA)
model <- lmer(Reaction ~ Days + (Days | Subject), dataA)
Then, I can use that model to generate predicted values from the second dataset (dataB)
pred_values <- predicted(model, data=dataB)
I can obtain the marginal R2 for the model applied to dataA using the MuMIn package
library(MuMIn)
r.squaredGLMM(model)
But how do I learn how well the model performed on dataB? I can correlate the predicted values with the actual values, and square it:
(cor.test(dataB$Reaction, pred_values)$estimate)^2
But this produces an overall R2, which is not ideal for interpretation purposes.