I am trying to use properly pipelines and column transformers from sklearn but always end up with an error. I reproduced it in the following example.
# Data to reproduce the error
X = pd.DataFrame([[1, 2 , 3, 1 ],
[1, '?', 2, 0 ],
[4, 5 , 6, '?']],
columns=['A', 'B', 'C', 'D'])
#SimpleImputer to change the values '?' with the mode
impute = SimpleImputer(missing_values='?', strategy='most_frequent')
#Simple one hot encoder
ohe = OneHotEncoder(handle_unknown='ignore', sparse=False)
col_transfo = ColumnTransformer(transformers=[
('missing_vals', impute, ['B', 'D']),
('one_hot', ohe, ['A', 'B'])],
remainder='passthrough'
)
Then calling the transformer as follows:
col_transfo.fit_transform(X)
Returns the following error:
TypeError: Encoders require their input to be uniformly strings or numbers. Got ['int', 'str']