I have a problem while trying to implement a pipeline, where I want to use the OrdinalEncoder and OneHotEncoder on different categorical columns.
At this point my code is as following:
X = stroke_df.drop(columns=['id', 'smoking_status', 'stroke'])
y = stroke_df['stroke'].copy()
num_columns = X.select_dtypes(np.number).columns.tolist()
cat_columns = X.select_dtypes('object').columns.tolist()
all_columns = num_columns + cat_columns # this order will need to be preserved
print('Numerical columns:', ', '.join(num_columns))
print('Categorical columns:', ', '.join(cat_columns))
num_pipeline = Pipeline([
('imputer', SimpleImputer(missing_values=np.nan, strategy='median')),
('scaler', StandardScaler())
])
cat_pipeline = ColumnTransformer([
('label_encoder', LabelEncoder(), ['ever_married', 'work_type']),
('one_hot_encoder', OneHotEncoder(), ['gender', 'residence_type'])
])
pipeline = ColumnTransformer([
('num', num_pipeline, num_columns),
('cat', cat_pipeline, cat_columns)
])
However after trying to call fit_transform on the pipeline and preprocess the input feature matrix I get TypeError:
X_prep = pipeline.fit_transform(X)
TypeError: fit_transform() takes 2 positional arguments but 3 were given