Python IndexError when using predict() function

Viewed 185

I'm trying to use Python's sklearn.naive_bayes CategoricalNB() model. I can train the model without errors, but when I try to predict I get the following error:

C:\ProgramData\Anaconda3\envs\Python\lib\site-packages\sklearn\naive_bayes.py in predict(self, X)
     81         check_is_fitted(self)
     82         X = self._check_X(X)
---> 83         jll = self._joint_log_likelihood(X)
     84         return self.classes_[np.argmax(jll, axis=1)]
     85 

C:\ProgramData\Anaconda3\envs\Python\lib\site-packages\sklearn\naive_bayes.py in _joint_log_likelihood(self, X)
   1459         for i in range(self.n_features_in_):
   1460             indices = X[:, i]
-> 1461             jll += self.feature_log_prob_[i][:, indices].T
   1462         total_ll = jll + self.class_log_prior_
   1463         return total_ll

IndexError: index 6 is out of bounds for axis 1 with size 6

I get the error on the last line of this code:

model = CategoricalNB()
model.fit(X_train, y_train)
y_train_pred = model.predict(X_train)
y_test_pred = model.predict(X_test) 

Where
X_train shape is (1318, 12) and type is numpy.ndarray
y_train shape is (1318,) and type is pandas.core.series.Series
X_test shape is (566, 12) and type is numpy.ndarray

The input variables underwent OrdinalEncoder() and for the target I used LabelEncoder(), so each column is some positive integer value between 0 and however many classes are in that variable. The target is multiclass with values 0 to 6.

I understand in principle what an Index error is, but I'm not sure why I'm getting one here since I'm not manually setting the range of indices anywhere. Any tips how I can troubleshoot or figure out what dimensions or datatypes I'm supposed to use? Am I setting up my dataframe incorrectly for a CategoricalNB model?

Thank you!


Update:

So I was able to solve this by adding stratify=y to the train_test_split() method as shown below.

from sklearn.model_selection import train_test_split

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size = 0.3, stratify=y, random_state = 123)

I'm still not sure why this solves an indexing problem, because I understand stratify to make sure there are a proportional amount of each class in the test and train splits.


Update 2:

My dataset has several output variables, and the above solution only works for some of them. I still get an IndexError for some of them.


Update 3:

So apparently this is a bug where a category shows in the test set that wasn't initially in the training set. See discussion on GitHub.

Does anyone have any ideas how to get around this bug?

Thank you!

0 Answers
Related