I have a large dataframe with 4 different variables ($name), that have many different values ($expl.val). Based on a model I made, I also have thirty separate predictions of ten different models, 3 runs for each ($pred.name). For each prediction I have a value (pred.value).
name expl.val pred.name pred.val
var1 0 RUN1_MODEL1 0.57727
var1 0 RUN2_MODEL1 0.561696
var1 0 RUN3_MODEL1 0.553354
var1 0 RUN1_MODEL2 0.719538
var1 0 RUN2_MODEL2 0.695912
var1 0 RUN3_MODEL2 0.729262
var2 5 RUN1_MODEL1 0.694463
var2 5 RUN2_MODEL1 0.699222
var2 5 RUN3_MODEL1 0.695147
var2 5 RUN1_MODEL4 0.886816
var2 5 RUN2_MODEL4 0.960639
var2 5 RUN3_MODEL4 0.982607
My ultimate goal is to plot response curves (x-axis: $expl.value, y-axis: $pred.value), but as the data frame is now, I would for each variable (%name) have 30 response curves, as I have 30 separate predictions per variable value. This would be messy.
Therefore, I'd like to eliminate the separate prediction runs per model, so that for each variable value I only have 10 predicted values (ten per model instead of thirty).
So the output would, in this case, look like this:
name expl.val pred.name pred.val
var1 0 RUNavg_MODEL1 0.564107
var1 0 RUNavg_MODEL2 0.714904
var2 5 RUNavg_MODEL1 0.696227
var2 5 RUNavg_MODEL4 0.943354
I am unsure how to approach this in Rstudio, because I want to average based on the different runs per different model, but then also separated by the value of the 4 variables on which the predictions depend. My only instinct is that I may want to split $pred.name so I get a new factor variable with 3 levels (RUN1,RUN2,RUN3) that I then somehow need to average not only by MODEL# but also by $expl.val.