In my multi-label classification problem apart from class predictions, I need to retrieve confidence scores for these predictions.
I'm using OneVsRestClassifer with a LogisticRegression model used as a base classifier. While experimenting with my training and test sets I noticed that when no probability calibration is used the majority of the confidence scores are within the range of 0.95 - 0.95 and, strangely, around 10% of scores are very close to zero (and then no labels are predicted by the classifier).
I've read though that LogisticRegression should already be well-calibrated, so could someone please explain why such a behaviour can be observed? I was expecting a more smooth distribution of the probabilities. Does it mean that for OneVsRestClassifier the good calibration of its logistic regression components no longer applies?
I decided to use the CalibratedClassifierCV class available in sklearn but I noticed that some probabilities significantly decreased as per below observations. How come a nearly hundred percent confidence can decrease to around 50%? Does anyone know any other method that could potentially help me scale those probabilities?
no calibration:
[0.99988209306050746], [0.99999511284844622], [0.99999995078223347], [0.99999989965720448], [0.99999986079273884], [0.99979651575446726], [0.99937347155943868]
isotonic calibration:
[0.49181127862298107], [0.62761741532720483], [0.71285392633212574], [0.74505221607398842], [0.67966429109225246], [0.47133458243199672], [0.48596255165026925]
sigmoid calibration:
[0.61111111111111116], [0.86111111111111116], [0.86111111111111116], [0.86111111111111116], [0.86111111111111116], [0.61111111111111116], [0.47222222222222227]
The code I'm using at the moment:
#Fit the classifier
clf = LogisticRegression(C=1., solver='lbfgs')
clf = CalibratedClassifierCV(clf, method='sigmoid')
clf = OneVsRestClassifier(clf)
mlb = MultiLabelBinarizer()
mlb = mlb.fit(train_labels)
train_labels = mlb.transform(train_labels)
clf.fit(train_profiles, train_labels)
#Predict probabilities:
probas = clf.predict_proba([x_test])