I am using XGBoostClassifier in an imbalanced scenario, so I want to include the sample_weight parameter into the fit() function. I am using Python and the scikit-learn API.
The problem is that it seems that XGBoost does not take into account the weight parameter. In addition, even if I say n_jobs=-1, the training is not parallelized at all (only 1 CPU is used).
This is the code I am using:
from xgboost import XGBClassifier
import copy
# Load data and split into X and y
train_data = pd.read_csv("my_path.csv")
y_train = train_data[TARGET_COL].to_list()
x_train = copy.deepcopy(train_data)
del x_train[TARGET_COL]
# Evaluate weights
w_yes, w_no = evaluate_weights(...)
sample_weights = []
for y in y_train:
if y == 1:
sample_weights.append(w_yes)
elif y == 0:
sample_weights.append(w_no)
model = XGBClassifier(n_jobs=-1)
# Train the model
trained_model = model.fit(x_train, y_train, sample_weight=sample_weights)
The w_yes values is 25 and w_no is 0.25.
Even if the weights are quite different (the original dataset is really imbalanced), it seems to produce no effects, such as the n_jobs parameter.
Any idea? Do I have to use directly the XGBoost syntax, without the sklearn API?