I'm trying to predict new data using my model (Gaussian Naive Bayes).
I can predict the X_test. But why when I create new random data to predict, it said 'X has a different shape than during fitting.' ?
I want to predict this df :
array([[25.0, '@gmail.com', 'Indonesia', 'Thursday', 'ideal', '0-1 year',
'1-2 years', '2-3 years', 1.5, 0.2, 0.0]], dtype=object)
The format is same with the X_test :
array([29.0, '@gmail.com', 'Indonesia', 'Thursday', 'ideal', '2-3 years',
'0-1 year', '0-1 year', 0.0, 0.016666666666666666, 0.0],
dtype=object).
But X_test has 11000+ rows.
When I tried to predict df :
model_joblib.predict(df)
It resulted like this :
---------------------------------------------------------------------------
ValueError Traceback (most recent call last)
<ipython-input-44-3bb7caad9f62> in <module>
----> 1 model_joblib.predict(df)
~\anaconda3\lib\site-packages\sklearn\utils\metaestimators.py in <lambda>(*args, **kwargs)
117
118 # lambda, but not partial, allows help() to work with update_wrapper
--> 119 out = lambda *args, **kwargs: self.fn(obj, *args, **kwargs)
120 # update the docstring of the returned function
121 update_wrapper(out, self.fn)
~\anaconda3\lib\site-packages\sklearn\pipeline.py in predict(self, X, **predict_params)
405 Xt = X
406 for _, name, transform in self._iter(with_final=False):
--> 407 Xt = transform.transform(Xt)
408 return self.steps[-1][-1].predict(Xt, **predict_params)
409
~\anaconda3\lib\site-packages\sklearn\feature_selection\_base.py in transform(self, X)
82 return np.empty(0).reshape((X.shape[0], 0))
83 if len(mask) != X.shape[1]:
---> 84 raise ValueError("X has a different shape than during fitting.")
85 return X[:, safe_mask(X, mask)]
86
ValueError: X has a different shape than during fitting.
When I tried to predict X_test :
model_joblib.predict(X_test),
it can predict and resulted :
array([0, 0, 0, ..., 1, 0, 0], dtype=int64)
Here is my joblib model :
Pipeline(steps=[('transformer',
ColumnTransformer(remainder='passthrough',
transformers=[('age_transformer',
FunctionTransformer(func=<function create_age_bins at 0x000002287B6BE4C0>),
['age']),
('onehotpipe',
OneHotEncoder(drop='first'),
['character_length', 'day',
'domain', 'country']),
('ordinarypipe',
OrdinalEncoder(cols=['last_time_open_email',
'last_time_ope...
'0-1 year': 1,
'1-2 years': 2,
'2-3 years': 3,
'more than 3 years': 4}}]),
['last_time_open_email',
'last_time_open_shopee',
'last_time_checkout_shopee']),
('robust', RobustScaler(),
['open_frequency',
'login_frequency',
'checkout_frequency'])])),
('selection', SelectPercentile(percentile=100)),
('resampling', SMOTE(k_neighbors=50, random_state=2021)),
('clf', GaussianNB())])
Can someone please explain me why I can not do this predict ?
Right now I'm forced to delete FunctionTransformer from my pipeline in order to make me able to predict the new dataset. But it cost me some metrics score.