ValueError: Class label 1 not present when specifying class_weight in RandomForestClassifier with k-fold cv

Viewed 133

I'm doing a binary classification on time-series data. Since it's for an academic project, I want to test classical ML models such as RandomForestClassifier as well.

However, while using TimeSeriesSplit K-fold Cross-Validation, it is possible that while training; labels have only one class instead of both, which is raising ValueError.

from sklearn.ensemble import RandomForestClassifier

rfc = RandomForestClassifier(class_weight={1:10, 0:1})
rfc.fit([[0, 0, 1], [1, 0, 1]], [0, 0])

This gives, ValueError: Class label 1 not present.

I know it doesn't make sense to train with only one label, but then it works fine if we don't specify class_weight. Is this a bug?

How do I get around this programmatically if I'm automating my testing?

1 Answers

I think the issue has been fixed on 11 Mar 2022.

I am able to run below code without any problem today:

from sklearn.ensemble import RandomForestClassifier

rfc = RandomForestClassifier(class_weight={1:10, 0:1})
rfc.fit([[0, 0, 1], [1, 0, 1]], [0, 1])
rfc.predict([[1, 1, 1]])

Output: array([1])


rfc.fit([[0, 0, 1], [1, 0, 1]], [0, 0])
rfc.predict([[1, 1, 1]])

Output: array([0])
Related