How to do Multi label classification or Multi class classification of the below problem? Pandas Python

Viewed 190

My original data looks like this.

    id  season   home_team  away_team  home_goals  away_goals result winner
 0  0   2006-07  Shu        Liv        1           1          D      NaN
 1  1   2006-07  Ars        Avl        1           1          D      NaN
 2  2   2006-07  Eve        Wat        2           1          H      Eve
 3  3   2006-07  New        Wig        2           1          H      New
 4  4   2006-07  Por        Bla        3           0          H      Por

The purpose is to build a model that predicts i.e.

Home Team Win 55%

Draw          13%

Away Team Win 32%

I Selected these 3 columns and label encoded them

home_team, away_team, winner

Then I created these new classes/lables.

df.loc[df["winner"]==df["home_team"],"home_team_win"]=1
df.loc[df["winner"]!=df["home_team"],"home_team_win"]=0

df.loc[df["result"]=='D',"draw"]=1
df.loc[df["result"]!='D',"draw"]=0

df.loc[df["winner"]==df["away_team"],"away_team_win"]=1
df.loc[df["winner"]!=df["away_team"],"away_team_win"]=0

Now the encoded data is looking like this,

    home_team   away_team   home_team_win   away_team_win   draw
0   28          19          0               0               1
1   1           2           0               0               1
2   14          34          1               0               0
3   23          37          1               0               0
4   25          4           1               0               0

Initially, I used the code below for a single label 'home_team_win' and it worked fine, but it doesn't support multi classes/labels.

X = prediction_df.drop(['home_team_win'] ,axis=1)

y = prediction_df['home_team_win']

logReg=LogisticRegression(solver='lbfgs')

rfe = RFE(logReg, 20)

rfe = rfe.fit(X, y.values.ravel())

How to do Multi label classification or Multi class classification of this problem?

1 Answers

The target binary variables home_team_win, away_team_win, and draw are mutually exclusive. It does not seem to be a good idea to use multi-label methods in this problem, since, in general, they are designed to exploit dependencies among labels, which is nonexistent in this dataset.

I suggest modelling it as a multi-class problem in its most common form, where there is a single column with three classes: 0,1, and 2 (representing home_team_loss, draw, away_team_win). Many implementations of classifiers in scikit-learn can work directly in this manner. Logistic Regression is one of them:

from sklearn.linear_model import LogisticRegression

logReg=LogisticRegression(solver='lbfgs', multi_class='ovr')
logReg.fit(X,Y)
logReg.predict_proba(X)

This code will output the desired probabilities for each class of each row of X. In particular, this code trains one Logistic Regression for each class separately (this is what the multi_class='ovr' parameter do).

Take a look at https://scikit-learn.org/stable/supervised_learning.html for other classifiers that directly work in this multi-class dataset form that I suggested.

Related