Confused about Column Order for sklearn Pipeline Imputer

Viewed 234

I have a pipeline that imputes and transforms my dataframe, but the resulting array seems to have the columns re-ordered. I'm struggling to understand how to setup my pipeline so that the order is retained.

# this ordering is the same as train.columns from when I fit the pipeline
cols = ['x1','x3','x5','x8']  # specific ordering

x = data[cols].copy()

imputer = mainmodel.named_steps['preprocessor'] #this is the imputation step that I had setup

x_transformed = imputer.transform(x[cols])

df_x_transformedx_transformed = pd.DataFrame(x_transformed, columns=cols)

When I look at the result, the column headers are mismatched to the columns. So something in the preprocessor pipeline is re-ordering the columns and I'm not sure how or where to control that.

I feel like this could be an area full of risk for a mistake, especially if I change [cols] to another order --- would imputer possibly accept any order and produce mistakes?

imputer

ColumnTransformer(remainder='passthrough',
                  transformers=[('num',
                                 Pipeline(steps=[('imputer',
                                                  SimpleImputer(fill_value=-9999,
                                                                strategy='constant'))]),
                                 ['x1']),
                                ('cat1',
                                 Pipeline(steps=[('imputer',
                                                  SimpleImputer(fill_value='missing',
                                                                strategy='constant')),
                                                 ('ord_encoding',
                                                  OrdinalEncoder())]),
                                 ['x4', 'x7',
                                  'x8', 'x11'])])
0 Answers
Related