I am using a very simple kaggle dataset to understand how SelectFromModel with a logistic regression works. The idea was to create a very simple pipeline with some basic data processing (dropping a column + scaling), pass it to feature selection (logreg) and then fit an xgboost model (not included in the code). By reading the documentation, my understanding is that given my X_train and y_train a logreg model is fitted and those features with coeffs higher or equal than the threshold are selected. In my case, I set the threshold to mean*1.25.
I cannot understand why the output selector.threshold_ is not the same as selector.estimator_.coef_.mean()*1.25. I was expecting to get the same value, why is it not the case?
Moving forward, I wanted to do GridSearchCV to fine-tune my pipeline parameters. I usually do it this way:
from sklearn.model_selection import GridSearchCV
params = {}
params['gradientboostingclassifier__learning_rate'] = [0.05, 0.1, 0.2]
params['selectfrommodel__estimator__C'] = [0.1, 1, 10]
params['selectfrommodel__estimator__penalty']= ['l1', 'l2']
params['selectfrommodel__estimator__threshold']=['median', 'mean', '1.25*mean', '0.75*mean']
grid = GridSearchCV(pipe, params, cv=5, scoring='recall')
%time grid.fit(X_train, y_train);
Unfortunately, it looks like the threshold is not in the list of parameters (pipe.named_steps.selectfrommodel.estimator.get_params().keys()) so for GridSearchCV to work this line needs to be commented.
params['selectfrommodel__estimator__threshold']=['median', 'mean', '1.25*mean', '0.75*mean']
Is there a way to fine-tune the threshold?
