Can tidymodels be used to implement the cross-validation scheme described in Henckaerts et al. (2020)?

Viewed 166

I am referring to the scheme described at page 11 of this paper: https://arxiv.org/pdf/1904.10890.pdf.

enter image description here

Instead of having a test set and large train set broken down into 5 folds, you take whole dataset and create 6 folds. Then each of the 6 folds is considered as the "test set" for the other 5 folds. The idea is the you end up using the whole dataset as a test set. Plus, you get 6 sets of performance metrics instead of just one.

I don't know much about tidymodels besides recipes (which I love). Will tidymodels allow me to do something similar to this, or should I just use {rsample} to create the 6 folds and then build a custom approach?

2 Answers

Can't comment yet, so have to post this as a full answer, sorry. Anyway: the tidymodels intro itself refers to {rsample} as part of the tidymodels framework. So I guess that's the package for the task. Then, putting each sample data into one row (by taking advantage of 'list columns', i.e. a column with a list / dataframe per cell) might help keep things compact. Example: Fit a different model for each row of a list-columns data frame

Yes, if I am understanding you correctly:

library(rsample)

car_split <- initial_split(mtcars)
cars_train <- training(car_split)
cars_test <- testing(car_split)

## CV folds for training set
vfold_cv(cars_train, v = 6)
#> #  6-fold cross-validation 
#> # A tibble: 6 × 2
#>   splits         id   
#>   <list>         <chr>
#> 1 <split [20/4]> Fold1
#> 2 <split [20/4]> Fold2
#> 3 <split [20/4]> Fold3
#> 4 <split [20/4]> Fold4
#> 5 <split [20/4]> Fold5
#> 6 <split [20/4]> Fold6


## CV folds for whole original data set
vfold_cv(mtcars, v = 6)
#> #  6-fold cross-validation 
#> # A tibble: 6 × 2
#>   splits         id   
#>   <list>         <chr>
#> 1 <split [26/6]> Fold1
#> 2 <split [26/6]> Fold2
#> 3 <split [27/5]> Fold3
#> 4 <split [27/5]> Fold4
#> 5 <split [27/5]> Fold5
#> 6 <split [27/5]> Fold6

Created on 2022-03-11 by the reprex package (v2.0.1)

You can create CV folds from any dataset you want, either a training set or an original whole data set. You might check out this chapter of our book for more on "spending your data budget".

Related