I am able to create train, validation, and test sets for one fold experiments using sklearn like below with train, val and test having a ratio of 60/20/20:
x_train, x_test, y_train, y_test = train_test_split(x, y, test_size=0.4, random_state=42)
x_test, x_val, y_test, y_val = train_test_split(x_test, y_test, test_size=0.5, random_state=42)
train = [x_train, y_train]
train_result = pd.concat(train, axis=1)
train_numpy = train_result.to_numpy()
np.savetxt('train_60.txt', train_numpy, fmt="%s", delimiter='\t')
test = [x_test, y_test]
test_result = pd.concat(test, axis=1)
test_numpy = test_result.to_numpy()
np.savetxt('test_20.txt', test_numpy, fmt="%s", delimiter='\t')
val = [x_val, y_val]
val_result = pd.concat(val, axis=1)
val_numpy = val_result.to_numpy()
np.savetxt('val_20.txt', val_numpy, fmt="%s", delimiter='\t')
I am also able to create 5 fold files for train and validation sets only using sklearn library (train 80/ val: 20):
kf = KFold(n_splits=5, random_state=42, shuffle=True)
for j, train_test in enumerate(kf.split(x, y)):
train_index, test_index = train_test
x_train, x_val = x[train_index], x[test_index]
y_train, y_val = y[train_index], y[test_index]
train = [x_train, y_train]
train_result = pd.concat(train, axis=1)
train_numpy = train_result.to_numpy()
np.random.shuffle(train_numpy)
train_name = 'train_fold' + str(j+1) + '.txt'
np.savetxt(train_name, train_numpy, fmt="%s", delimiter='\t')
val = [x_val, y_val]
val_result = pd.concat(val, axis=1)
val_numpy = val_result.to_numpy()
np.random.shuffle(val_numpy)
val_name = 'val_fold' + str(j+1) + '.txt'
np.savetxt(val_name, val_numpy, fmt="%s", delimiter='\t')
So, I have two questions, how can I automatically produce 5 fold train, test, val txt files like above? How can these two methods basically be combined and second question is that does the ratio of 60/20/20 make sense for train/test/val in 5-fold experiment or should we definitely choose to go with 80/10/10 ratio?
One basic idea which is rather manual is to do 5 fold CV and then take the first half of each val_foldi.txt as val and the next half as test. However, this is not automatic and also not sure if it is fully acceptable by the ML community. Also, this approach only would work for 80/10/10 split.