I have a pipeline that imputes and transforms my dataframe, but the resulting array seems to have the columns re-ordered. I'm struggling to understand how to setup my pipeline so that the order is retained.
# this ordering is the same as train.columns from when I fit the pipeline
cols = ['x1','x3','x5','x8'] # specific ordering
x = data[cols].copy()
imputer = mainmodel.named_steps['preprocessor'] #this is the imputation step that I had setup
x_transformed = imputer.transform(x[cols])
df_x_transformedx_transformed = pd.DataFrame(x_transformed, columns=cols)
When I look at the result, the column headers are mismatched to the columns. So something in the preprocessor pipeline is re-ordering the columns and I'm not sure how or where to control that.
I feel like this could be an area full of risk for a mistake, especially if I change [cols] to another order --- would imputer possibly accept any order and produce mistakes?
imputer
ColumnTransformer(remainder='passthrough',
transformers=[('num',
Pipeline(steps=[('imputer',
SimpleImputer(fill_value=-9999,
strategy='constant'))]),
['x1']),
('cat1',
Pipeline(steps=[('imputer',
SimpleImputer(fill_value='missing',
strategy='constant')),
('ord_encoding',
OrdinalEncoder())]),
['x4', 'x7',
'x8', 'x11'])])