Sklearn Pipeline: Get final list of feature names after ColumnTransformer followed by SelectFromModel from Pipeline

Viewed 187

I am using the following Pipeline:

numeric_transformer = Pipeline(steps = [('numeric_missing', MeanMedianImputer(imputation_method="median")),
('numeric_scale', StandardScaler())])

categorical_transformer = Pipeline(steps=[('categorical_missing', CategoricalImputer(imputation_method='missing')),
  ('ordinal', OrdinalEncoder(encoding_method='arbitrary'))])

preprocessor = ColumnTransformer(
transformers=[('numeric_transformation', numeric_transformer, numeric_features_list),
('categorical_transformation', categorical_transformer, categorical_features_list)],
remainder = 'passthrough')

pipe = Pipeline(steps=[('preprocessor', preprocessor),
                      ('feature_selection', SelectFromModel(XGBRegressor(random_state = 1,warm_start = False, silent=True, verbosity=0))), 
                      ('regressor', XGBRegressor(random_state = 1,
                      warm_start = False, silent=True, verbosity=0))
                      ])

search_space = [{'regressor__loss': ['ls', 'lad', 'huber', 'quantile'],
            'regressor__learning_rate': [1e-3, 1],
            'regressor__n_estimators': [25, 500],
            'regressor__subsample': [0.01, 0.99],
            'regressor__max_depth':[1, 10],
            'regressor__max_features':['log2','sqrt'],
            'regressor__alpha':[1e-5, 1e-2, 0.1, 1, 100],
            'regressor__lambda':[1e-5, 1e-2, 0.1, 1, 100]
            }]

clf = RandomizedSearchCV(pipe, search_space, verbose=1, 
                        scoring='neg_root_mean_squared_error', cv=5, 
                        refit=True, random_state=1, n_iter=5)

clf = clf.fit(X_train, Y_train) 

p=pipe[:-1].fit(X_train,Y_train)

I want to replicate the results i.e. get the 'preprocessor' and 'feature_selection' steps done on my training and testing dataset separately so as to use it in other functions. Im using the following codes:

X_train_tf=pd.DataFrame(p.transform(X_train), **columns=___**)

X_test_tf=pd.DataFrame(p.transform(X_test), **columns =___**)

How should I get the feature names to use in columns in the above code from the Pipeline that gives me features as an output from the 'preprocessor' and 'feature_selection' steps?

Note: I am coding in Python on Databricks.

0 Answers
Related