sk-learn pipeline with heterogeneous data: apply PCA to continous features only, after ColumnTransform

Viewed 13

I have been trying to do something which I think should be possible but cannot work it out.

At the moment I have a pipeline which looks like this: Pipeline diagram

As you can see from the diagram the PCA is applied after all the ColumnsTranforms, on all tranformed features.

I would like it to apply only to the continuous features. Where it breaks down is that ColumnTransformer outputs numpy arrays and so I cannot refer back to the column names. I tried and get the ubiquitous error ValueError: Specifying the columns using strings is only supported for pandas DataFrames

Here is the code to create the pipeline.

from sklearn.linear_model import LogisticRegression
from sklearn import set_config
from sklearn.compose import make_column_transformer

set_config(display="diagram")

# define the preprocessing steps. We must take care not to preprocess  features twice.
# define the ordinal encoder
ordi = OrdinalEncoder(categories=[[0,1]])
# define the OneHotEncoder
cat =OneHotEncoder(sparse = False)

preproc = ColumnTransformer([
  ('log_trans',log_scale,log_features),
  ('scaler',cont,continuous),
  ('long',long_scale,long_features),
  ('lat',lat_scale,lat_features),
  ('reg',ordi,discrete),
  ('deciles',KBin,kb_features),
  ('categ',cat,categorical)
],remainder = 'passthrough')


logreg = make_pipeline(
  preproc,
  pca_full,
  LogisticRegression(C =100,multi_class = 'multinomial',solver ='saga',max_iter = 4000,tol 
  =0.001))
logreg

continuous feature names are stored in the lists continuous , log_features,long_features and lat_features

There have been a few similar posts in the past but they go back some years so I am hopeful someone might have an up to date answer. previous stackoverflow question

0 Answers
Related