I have been trying to do something which I think should be possible but cannot work it out.
At the moment I have a pipeline which looks like this: Pipeline diagram
As you can see from the diagram the PCA is applied after all the ColumnsTranforms, on all tranformed features.
I would like it to apply only to the continuous features. Where it breaks down is that ColumnTransformer outputs numpy arrays and so I cannot refer back to the column names.
I tried and get the ubiquitous error ValueError: Specifying the columns using strings is only supported for pandas DataFrames
Here is the code to create the pipeline.
from sklearn.linear_model import LogisticRegression
from sklearn import set_config
from sklearn.compose import make_column_transformer
set_config(display="diagram")
# define the preprocessing steps. We must take care not to preprocess features twice.
# define the ordinal encoder
ordi = OrdinalEncoder(categories=[[0,1]])
# define the OneHotEncoder
cat =OneHotEncoder(sparse = False)
preproc = ColumnTransformer([
('log_trans',log_scale,log_features),
('scaler',cont,continuous),
('long',long_scale,long_features),
('lat',lat_scale,lat_features),
('reg',ordi,discrete),
('deciles',KBin,kb_features),
('categ',cat,categorical)
],remainder = 'passthrough')
logreg = make_pipeline(
preproc,
pca_full,
LogisticRegression(C =100,multi_class = 'multinomial',solver ='saga',max_iter = 4000,tol
=0.001))
logreg
continuous feature names are stored in the lists continuous , log_features,long_features and lat_features
There have been a few similar posts in the past but they go back some years so I am hopeful someone might have an up to date answer. previous stackoverflow question