I am still new to this and have a lot to learn, so bear with me while I learn how to do this well. I took a Kaggle data set and wanted to work on using pipelines to make my code neater and to simplify the process of cleaning up my categorical and numerical data. I have had no errors in my work until I tried fitting my model at the end, where it gave me ValueError: too many values to unpack (expected 3). Any help would be greatly appreciated!
I have looked through other questions similar to mine, but haven't found one that related specifically to the issue I am having. I thought it might have something to do with my naming of variables, but so far I haven't been able to discover my error.
df = pd.read_csv('master.csv')
df = df.drop('country-year', axis=1)
df[' gdp_for_year ($) '] = df[' gdp_for_year ($) '].str.replace(',','')
df[' gdp_for_year ($) '] = df[' gdp_for_year ($) '].astype(str).astype(float)
#print(df.info())
y = df.suicides_no
features = ['country', 'sex', 'age', 'generation', 'year', 'population',
'HDI for year', ' gdp_for_year ($) ', 'gdp_per_capita ($)']
X = df[features].copy()
X_train, X_valid, y_train, y_valid = train_test_split(X, y, train_size=0.8, test_size=0.2,random_state=0)
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import OneHotEncoder
numerical_transformer = SimpleImputer(strategy='mean')
categorical_transformer = Pipeline(steps=[
('imputer', SimpleImputer(strategy='most_frequent')),
('onehot', OneHotEncoder(handle_unknown='ignore'))
])
preprocessor = ColumnTransformer(
transformers=[
('num', numerical_transformer, 'suicides_no', 'population', 'HDI for year', ' gdp_for_year ($) ', 'gdp_per_capita ($)'),
('cat', categorical_transformer, 'country', 'sex', 'generation', 'year', 'age')
])
model_1 = RandomForestRegressor(n_estimators=100, random_state=0)
my_pipeline = Pipeline(steps=[('preprocessor', preprocessor),
('model', model_1)
])
my_pipeline.fit(X_train, y_train)
Traceback (most recent call last):
File "...Suicide Rates/SR.py", line 86, in <module>
my_pipeline.fit(X_train, y_train)
File "...pipeline.py", line 265, in fit
Xt, fit_params = self._fit(X, y, **fit_params)
File "...pipeline.py", line 230, in _fit
**fit_params_steps[name])
File "...memory.py", line 342, in __call__
return self.func(*args, **kwargs)
File "...pipeline.py", line 614, in _fit_transform_one
res = transformer.fit_transform(X, y, **fit_params)
File "...compose\_column_transformer.py", line 445, in fit_transform
self._validate_transformers()
File "...ompose\_column_transformer.py", line 256, in _validate_transformers
names, transformers, _ = zip(*self.transformers)
ValueError: too many values to unpack (expected 3)