How do I know that RepeatedStratifiedKFold is actually working as it should?

Viewed 2105

I am trying to choose a method to split my data into train and test sets. I am currently using Scikit's RepeatedStratifiedKFold. According to the documentation, the RepeatedStratifiedKFold is a:

Repeated Stratified K-Fold cross validator.

Repeats Stratified K-Fold n times with different randomization in each repetition.

I use the RepeatedStratifiedKFold using 5 folds and 100 repetitions on a dataset consisting of 1000 observations as follows:

rskf = RepeatedStratifiedKFold(n_splits=5, n_repeats=100, random_state=None)

for train_index, test_index in rskf.split(X, y):

   X_train, _X_test = X[train_index], X[test_index]

   y_train, y_test = y[train_index], y[test_index]

However, when I look at the X_train set, I only see 800 observations (4 train folds). Shouldn't it contain all 100 train sets as per the number of repetitions?

My second question: after splitting your data using the RepeatedStratifiedKFold method, what happens when you fit your classification model on the X_train and y_train datasets? Does the model train on all 100 repetitions?

Suppose I just wanted the F1-score from the model after testing it. Does it give me the average score across all 100 repetitions?

Thanks!

1 Answers

However, when I look at the X_train set, I only see 800 observations (4 train folds). Shouldn't it contain all 100 train sets as per the number of repetitions?

Not really. Actually, when you loop through rksf.split(X, y), you are looping through the number of iterations (given by n_splits * n_repeats). In this case, you are looping through 5 * 100 = 500 iterations - each of them with different partitions. You can very easily check this by adding a counter to your loop and printing it:

ii = 1
for train_index, test_index in rskf.split(X, y):
   print(f"Iteration {ii}")

   X_train, _X_test = X[train_index], X[test_index]
   y_train, y_test = y[train_index], y[test_index]

   ii += 1

Which would result in

Iteration 1
Iteration 2
...
Iteration 499
Iteration 500

Thus, it makes sense that when you run your code you see that X_train has 800 observations. That corresponds to the partition of one of your iterations. If this isn't clear yet, I suggest you take a look at this other SO answer.


My second question: after splitting your data using the RepeatedStratifiedKFold method, what happens when you fit your classification model on the X_train and y_train datasets? Does the model train on all 100 repetitions?

No, the way you have it currently defined, the model would be fitted using X_train and y_train of a single iteration.


Suppose I just wanted the F1-score from the model after testing it. Does it give me the average score across all 100 repetitions?

No, related to your earlier question, you would get the F1-score of a single iteration. You could save the F1-score of each iteration in a list and then get calculate the average (and maybe also SD). You could also consider using Scikit's cross_val_score

Related