For hyperparameter tuning, I use the function GridSearchCV from the Python package sklearn. Some of the models that I test require feature scaling (e.g. Support Vector Regression - SVR). Recently, in the Udemy course Machine Learning A-Z™: Hands-On Python & R In Data Science, the instructors mentioned that for SVR, the target should also be scaled (if it is not binary). Bearing this in mind, I wonder whether the target is also scaled in each iteration of the cross-validation procedure performed by GridSearchCV or if only the features are scaled. Please see the code below, which illustrates the normal procedure that I use for hyperparameter tuning for estimators that require scaling for the training sets:
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import GridSearchCV
from sklearn.svm import SVR
def SVRegressor(**kwargs):
'''contruct a pipeline to perform SVR regression'''
return make_pipeline(StandardScaler(), SVR(**kwargs))
params = {'svr__kernel': ["poly", "rbf"]}
grid_search = GridSearchCV(SVRegressor(), params)
grid_search.fit(X, y)
I know that I could simply scale X and y a priori and drop the StandardScaler from the pipeline. However, I want to implement this approach in a code pipeline where multiple models are tested, in which, some require scaling and others do not. That is why I want to know how GridSearchCV handles scaling under the hood.