The CatBoost documentation says the randomized_search method can accept train and test splits via the cv parameter, instead of defining a cross validation approach. To do this, one should provide:
An iterable yielding train and test splits as arrays of indices.
How do we define this object?
As a broken example, say my feature dataset has 10 rows. I want to use the first 5 rows for training, and the last 5 rows for validation/testing.
I extract the index values
train_index = X[0:5].index
test_index = X[5:10].index
I supply the indexes to the randomized_search method
a_search = model.randomized_search(param_distributions=params,
X = X,
y = y,
n_iter=5,
cv={train_index,test_index})
This set that I provide in cv={train_index,test_index} is a non-starter, as it's not iterable, but I am at a loss as to how such an iterable should look. I simply want to define which rows of X and y should be used for training, and which for testing. The goal is to speed up training by dispensing with cross validation, and using a dedicated validation dataset.