Hyperparameter optimisation in Python with a separate validation set

Viewed 118

I am trying to optimise the hyper parameters of a random forest regressor in Python.

I have 3 separate datasets: train/validate/test. Therefore, rather than using a cross validation method I want to use the specific validation set to tune the hyperparameters, i.e. the "First Approach" described in this stackoverflow post.

Now, sklearn has some nice inbuilt methods for hyperparameter optimisation using cross validation (e.g. this tutorial), but what about if I want to tune my hyperparameters with a specific validation set? Is it still possible to use a method like RandomizedSearchCV?

2 Answers

It is indeed possible with cv option. As the documentation suggests, one of the possible inputs is an iterable of train/test index tuples:

An iterable yielding (train, test) splits as arrays of indices.

So, a list of size one with train and validation indices packed as a tuple would be ok.

I think we should just have some wording clarified:

'Validation set'

A validation-set is used to evaluate your model on a unseen set of data i.e data not used for training. This is to simulate how your model would behave on new data. We use the validation-set to tune our hyper-parameters such as number of trees, max-depths etc. and chose the hyper-parameters which works best on the validation set.

'Cross-validate'

When you CV (cross-validate) with, say, 5 folds you divide your data into 5 sets where set [1,2,3,4] are used for traning, and set 5 is used for validation. Then you use [2,3,4,5] for training and use set 1 for validation - you repeat this untill all sets (i.e 5 times when using 5 fold) have been used as a validation-set and then you would average your 5 validation-score e.g accuracy to get one score which you want to (often) maximize.

Answer

So, to answer your question; yes, you can use GridSearchCV on your validation-set but that wouldn't often be the case since. You would often do one of the following:

a) Use a (i.e one) validation-set to tune your hyper-parameters against, as explained in "Validation set"

b) Use all your data i.e train+validation as one data-set and then run a, say, 5-fold grid-CV search as explained in "Cross-validate"

Related