I am using the following Pipeline:
numeric_transformer = Pipeline(steps = [('numeric_missing', MeanMedianImputer(imputation_method="median")),
('numeric_scale', StandardScaler())])
categorical_transformer = Pipeline(steps=[('categorical_missing', CategoricalImputer(imputation_method='missing')),
('ordinal', OrdinalEncoder(encoding_method='arbitrary'))])
preprocessor = ColumnTransformer(
transformers=[('numeric_transformation', numeric_transformer, numeric_features_list),
('categorical_transformation', categorical_transformer, categorical_features_list)],
remainder = 'passthrough')
pipe = Pipeline(steps=[('preprocessor', preprocessor),
('feature_selection', SelectFromModel(XGBRegressor(random_state = 1,warm_start = False, silent=True, verbosity=0))),
('regressor', XGBRegressor(random_state = 1,
warm_start = False, silent=True, verbosity=0))
])
search_space = [{'regressor__loss': ['ls', 'lad', 'huber', 'quantile'],
'regressor__learning_rate': [1e-3, 1],
'regressor__n_estimators': [25, 500],
'regressor__subsample': [0.01, 0.99],
'regressor__max_depth':[1, 10],
'regressor__max_features':['log2','sqrt'],
'regressor__alpha':[1e-5, 1e-2, 0.1, 1, 100],
'regressor__lambda':[1e-5, 1e-2, 0.1, 1, 100]
}]
clf = RandomizedSearchCV(pipe, search_space, verbose=1,
scoring='neg_root_mean_squared_error', cv=5,
refit=True, random_state=1, n_iter=5)
clf = clf.fit(X_train, Y_train)
p=pipe[:-1].fit(X_train,Y_train)
I want to replicate the results i.e. get the 'preprocessor' and 'feature_selection' steps done on my training and testing dataset separately so as to use it in other functions. Im using the following codes:
X_train_tf=pd.DataFrame(p.transform(X_train), **columns=___**)
X_test_tf=pd.DataFrame(p.transform(X_test), **columns =___**)
How should I get the feature names to use in columns in the above code from the Pipeline that gives me features as an output from the 'preprocessor' and 'feature_selection' steps?
Note: I am coding in Python on Databricks.