How does XGBClassifier identify positive labels for scale_pos_weight parameter?

Viewed 198

I'm currently setting up a model using XGBClassifier() for a binary classification task in an unbalanced dataset (positive:negative = 1:20). Positive label is encoded as 1, negative label is encoded as 0. When trying to optimize the parameter scale_pos_weight via GridSearchCV(scoring=roc_auc) with otherwise default parameters, my assumption was that values >1 and roughly around 20 would perform better, as a typical value per documentation is sum(negative instances) / sum(positive instances). However, the gridsearch always ends up prefering the lowest value in the grid as low as scale_pos_weight=0.1.

My guess is that the model considers 0 as positive class. Is there any way to confirm/change that? The documentation I found offers no explanation on how XGBClassifier() identifies the positive label.

As a sidenote, performing the same gridsearch using catboost.CatBoostClassifier() with default settings results in higher optimal values for scale_pos_weight as expected. Thank you in advance!

0 Answers
Related