I'm trying to use Python's sklearn.naive_bayes CategoricalNB() model. I can train the model without errors, but when I try to predict I get the following error:
C:\ProgramData\Anaconda3\envs\Python\lib\site-packages\sklearn\naive_bayes.py in predict(self, X)
81 check_is_fitted(self)
82 X = self._check_X(X)
---> 83 jll = self._joint_log_likelihood(X)
84 return self.classes_[np.argmax(jll, axis=1)]
85
C:\ProgramData\Anaconda3\envs\Python\lib\site-packages\sklearn\naive_bayes.py in _joint_log_likelihood(self, X)
1459 for i in range(self.n_features_in_):
1460 indices = X[:, i]
-> 1461 jll += self.feature_log_prob_[i][:, indices].T
1462 total_ll = jll + self.class_log_prior_
1463 return total_ll
IndexError: index 6 is out of bounds for axis 1 with size 6
I get the error on the last line of this code:
model = CategoricalNB()
model.fit(X_train, y_train)
y_train_pred = model.predict(X_train)
y_test_pred = model.predict(X_test)
Where
X_train shape is (1318, 12) and type is numpy.ndarray
y_train shape is (1318,) and type is pandas.core.series.Series
X_test shape is (566, 12) and type is numpy.ndarray
The input variables underwent OrdinalEncoder() and for the target I used LabelEncoder(), so each column is some positive integer value between 0 and however many classes are in that variable. The target is multiclass with values 0 to 6.
I understand in principle what an Index error is, but I'm not sure why I'm getting one here since I'm not manually setting the range of indices anywhere. Any tips how I can troubleshoot or figure out what dimensions or datatypes I'm supposed to use? Am I setting up my dataframe incorrectly for a CategoricalNB model?
Thank you!
Update:
So I was able to solve this by adding stratify=y to the train_test_split() method as shown below.
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size = 0.3, stratify=y, random_state = 123)
I'm still not sure why this solves an indexing problem, because I understand stratify to make sure there are a proportional amount of each class in the test and train splits.
Update 2:
My dataset has several output variables, and the above solution only works for some of them. I still get an IndexError for some of them.
Update 3:
So apparently this is a bug where a category shows in the test set that wasn't initially in the training set. See discussion on GitHub.
Does anyone have any ideas how to get around this bug?
Thank you!