Dear colleagues I have created an scikit learn pipeline to traing and tune different HistBoostRegressors.
from scipy.stats import loguniform
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.feature_selection import VarianceThreshold
from sklearn.multioutput import MultiOutputRegressor
from sklearn.model_selection import RandomizedSearchCV
class loguniform_int:
"""Integer valued version of the log-uniform distribution"""
def __init__(self, a, b):
self._distribution = loguniform(a, b)
def rvs(self, *args, **kwargs):
"""Random variable sample"""
return self._distribution.rvs(*args, **kwargs).astype(int)
data_train, data_test, target_train, target_test = train_test_split(
df.drop(columns=TARGETS),
df[target_dict],
random_state=42)
pipeline_hist_boost_mimo_inside = Pipeline([('scaler', StandardScaler()),
('variance_selector', VarianceThreshold(threshold=0.03)),
('estimator', MultiOutputRegressor(HistGradientBoostingRegressor(loss='poisson')))])
parameters = {
'estimator__estimator__l2_regularization': loguniform(1e-6, 1e3),
'estimator__estimator__learning_rate': loguniform(0.001, 10),
'estimator__estimator__max_leaf_nodes': loguniform_int(2, 256),
'estimator__estimator__max_leaf_nodes': loguniform_int(2, 256),
'estimator__estimator__min_samples_leaf': loguniform_int(1, 100),
'estimator__estimator__max_bins': loguniform_int(2, 255),
}
random_grid_inside = RandomizedSearchCV(estimator=pipeline_hist_boost_mimo_inside, param_distributions=parameters, random_state=0, n_iter=50,
n_jobs=-1, refit=True, cv=3, verbose=True,
pre_dispatch='2*n_jobs',
return_train_score=True)
results_inside_train = random_grid_inside.fit(data_train, target_train)
However now I would like to know if it would be possible to pass different feature names to the step pipeline_hist_boost_mimo_inside["estimator"].
I have noticed that in the documentation of the multi output regressor we have a parameter call feature_names:
feature_names_in_ndarray of shape (n_features_in_,) Names of features seen during fit. Only defined if the underlying estimators expose such an attribute when fit.
New in version 1.0.
I have also found some documentation in scikit learn column selector which has the argument:
patternstr, default=None Name of columns containing this regex pattern will be included. If None, column selection will not be selected based on pattern.
The problem is that this pattern will depend on the target that I am fitting.
Is there a way to do this elegantly?
EDIT: Example of the dataset:
feat1, feat2, feat3.... target1, target2, target3....
1 47 0.65 0 0.5 0.6
The multioutput regressor will fit an histogram regressor for every pair of (feat1, feat2, feat3 and targetn). In the example of the table below I will have a pipeline which estimator step will contain a list of 3 estimators as a have 3 targets.
The question is how to pass for instance feat1 and feat2 to target1 but pass feat1 and feat3 to target2.