I'm trying to use Column Transformer with OneHotEncoder to transform my categorical data :
A quick look at my data :
I want to do one-hot-encoding for 3 features : 'sex' , 'smoker' , 'region', so I use Column Transformer by scikit-learn. ( I don't want to want to seperate numerical one and categorical one than transform them seperately, I just want to perform them on a single dataset)
My code :
cat_feature = X.select_dtypes(include = 'object') #select only categorical columns
enc = ColumnTransformer([ ('one_hot_encoder' , OneHotEncoder() , cat_feature ) ] ,
remainder = 'passthrough')
X_transformed = enc.fit_transform(X) # transformed version of original data
My problem is that, X_transformed is then removed all the feature names which is little bit confusing for me to debug :
So is there anyway to retain my columns' names after doing this transformation? I want to incorporate this transformer into a pipeline so I can't use pd.get_dummies.
Thank you!!

