According to the SVC documentation:
Platt scaling uses the logistic function
1 / (1 + exp(decision_value * probA_ + probB_))whereprobA_andprobB_are learned from the dataset.
now, if I try a simple SVC example
import numpy as np
from sklearn.datasets import load_digits
from sklearn.svm import SVC
X, y = load_digits(return_X_y=True)
mask = (y == 0) | (y == 1)
X = X[mask, :]
y = y[mask] # make a binary classification problem
model = SVC(probability=True)
model.fit(X,y)
If I use predict_proba to get all probabilities, this leads to
pd.DataFrame(model.predict_proba(X)).iloc[:,1]
0 0.00143866
1 0.99998592
2 0.00469432
3 0.99741052
4 0.00264468
while using the formula from above leads to a different result
pd.DataFrame(1/(1+(np.exp(model.decision_function(X) * model.probA_ + model.probB_)))).iloc[:,0]
0 0.00116055
1 0.99732754
2 0.00377729
3 0.99679562
4 0.00213140
Differences get much bigger than that in real world examples.
So what am I missing here? How does the formula need to be changed to make the results match with predict_proba?