'ValueError: X has a different shape than during fitting' when trying to predict new dataset

Viewed 176

I'm trying to predict new data using my model (Gaussian Naive Bayes).

I can predict the X_test. But why when I create new random data to predict, it said 'X has a different shape than during fitting.' ?

I want to predict this df :

array([[25.0, '@gmail.com', 'Indonesia', 'Thursday', 'ideal', '0-1 year',
                '1-2 years', '2-3 years', 1.5, 0.2, 0.0]], dtype=object)
    

The format is same with the X_test :

array([29.0, '@gmail.com', 'Indonesia', 'Thursday', 'ideal', '2-3 years',
               '0-1 year', '0-1 year', 0.0, 0.016666666666666666, 0.0],
              dtype=object). 
    

But X_test has 11000+ rows.

When I tried to predict df :

model_joblib.predict(df)

It resulted like this :

---------------------------------------------------------------------------
ValueError                                Traceback (most recent call last)
<ipython-input-44-3bb7caad9f62> in <module>
----> 1 model_joblib.predict(df)

~\anaconda3\lib\site-packages\sklearn\utils\metaestimators.py in <lambda>(*args, **kwargs)
    117 
    118         # lambda, but not partial, allows help() to work with update_wrapper
--> 119         out = lambda *args, **kwargs: self.fn(obj, *args, **kwargs)
    120         # update the docstring of the returned function
    121         update_wrapper(out, self.fn)

~\anaconda3\lib\site-packages\sklearn\pipeline.py in predict(self, X, **predict_params)
    405         Xt = X
    406         for _, name, transform in self._iter(with_final=False):
--> 407             Xt = transform.transform(Xt)
    408         return self.steps[-1][-1].predict(Xt, **predict_params)
    409 

~\anaconda3\lib\site-packages\sklearn\feature_selection\_base.py in transform(self, X)
     82             return np.empty(0).reshape((X.shape[0], 0))
     83         if len(mask) != X.shape[1]:
---> 84             raise ValueError("X has a different shape than during fitting.")
     85         return X[:, safe_mask(X, mask)]
     86 

ValueError: X has a different shape than during fitting.
    

When I tried to predict X_test :

model_joblib.predict(X_test),

it can predict and resulted :

array([0, 0, 0, ..., 1, 0, 0], dtype=int64)

Here is my joblib model :

Pipeline(steps=[('transformer',
                     ColumnTransformer(remainder='passthrough',
                                       transformers=[('age_transformer',
                                                      FunctionTransformer(func=<function create_age_bins at 0x000002287B6BE4C0>),
                                                      ['age']),
                                                     ('onehotpipe',
                                                      OneHotEncoder(drop='first'),
                                                      ['character_length', 'day',
                                                       'domain', 'country']),
                                                     ('ordinarypipe',
                                                      OrdinalEncoder(cols=['last_time_open_email',
                                                                           'last_time_ope...
                                                                                           '0-1 year': 1,
                                                                                           '1-2 years': 2,
                                                                                           '2-3 years': 3,
                                                                                           'more than 3 years': 4}}]),
                                                      ['last_time_open_email',
                                                       'last_time_open_shopee',
                                                       'last_time_checkout_shopee']),
                                                     ('robust', RobustScaler(),
                                                      ['open_frequency',
                                                       'login_frequency',
                                                       'checkout_frequency'])])),
                    ('selection', SelectPercentile(percentile=100)),
                    ('resampling', SMOTE(k_neighbors=50, random_state=2021)),
                    ('clf', GaussianNB())])
    
    

Can someone please explain me why I can not do this predict ?

Right now I'm forced to delete FunctionTransformer from my pipeline in order to make me able to predict the new dataset. But it cost me some metrics score.

0 Answers
Related