Getting class probabilities p(c|x) for each data point using sklearn

Viewed 133

How can we get the class probability of each test data point? There are some classifiers that do contain "predict_proba()" function that returns the probability of the data class the data point belongs to.

But there is no such function defined in the https://sklearn-lvq.readthedocs.io/en/stable/rslvq.html#

I need to calculate the class probabilities of each class so that the reject option can be applied. The idea is to calculate p(c|x) and check if the value is less than the threshold, then the data point should be rejected.

1 Answers

You can try this: Method 1: overriding the predict function as follows

from sklearn_lvq import RslvqModel
from sklearn.utils.validation import check_is_fitted
from sklearn.utils import validation
import numpy as np

class RslvqModel_custom(RslvqModel):

    def predict(self, x):
        """Predict class membership index for each input sample.

        This function does classification on an array of
        test vectors X.


        Parameters
        ----------
        x : array-like, shape = [n_samples, n_features]


        Returns
        -------
        C : array, shape = (n_samples,)
            Returns predicted values.
        """
        check_is_fitted(self, ['w_', 'c_w_'])
        x = validation.check_array(x)
        if x.shape[1] != self.w_.shape[1]:
            raise ValueError("X has wrong number of features\n"
                             "found=%d\n"
                             "expected=%d" % (self.w_.shape[1], x.shape[1]))

        def foo(e):
            fun = np.vectorize(lambda w: self._costf(e, w),
                               signature='(n)->()')
            pred = fun(self.w_)
            return pred

        predictions = np.vectorize(foo, signature='(n)->(n)')(x)
        sum = np.sum(predictions, axis=1).reshape(predictions.shape[0], 1)
        return predictions/sum

np.random.seed(1)
print(__doc__)
nb_ppc = 100
x = np.append(
    np.random.multivariate_normal([0, 0], np.eye(2) / 2, size=nb_ppc),
    np.random.multivariate_normal([5, 0], np.eye(2) / 2, size=nb_ppc), axis=0)
y = np.append(np.zeros(nb_ppc), np.ones(nb_ppc), axis=0)

rslvq = RslvqModel_custom(initial_prototypes=[[5,0,0],[0,0,1]]) #_custom
model = rslvq.fit(x, y)
predictions = model.predict([[3.67, 6.50], [4.97, 1.49], [1.14, -4.3]])

print('============================================================')
print('Predictions: ', predictions)
print('-------------------------------------------------------------')

OUTPUT:

Predictions:  [[0.5977081  0.4022919 ]
 [0.945568   0.054432  ]
 [0.33533978 0.66466022]]

Method 2: Or you can simply make use of builtin posterior method:

for data in [[3.67, 6.50], [4.97, 1.49], [1.14, -4.3]]:
    print('Class 0:  posterior: ', model.posterior(0, data))
    print('Class 1:  posterior: ', model.posterior(1, data))
    print('='*100)

OUTPUT:

-------------------------------------------------------------
Class 0:  posterior:  [[0.5977081]]
Class 1:  posterior:  [[0.4022919]]
========================================================
Class 0:  posterior:  [[0.945568]]
Class 1:  posterior:  [[0.054432]]
========================================================
Class 0:  posterior:  [[0.33533978]]
Class 1:  posterior:  [[0.66466022]]
========================================================
Related