I am new to Machine Learning. In a binary classfication problem, we encode/transform target variable like yes=1 and No=0 (directly in dataset) it gives follwoing results
- Accuracy:95
- Recall: 90
- Precision:94
- F1: 92
but if we encode/transform target variable inversely like yes=0 and No=1(directly in dataset), then it gives these results
- Accuracy:95
- Recall:97
- Precision:94
- F1:95
I am using XGboost algorithm. All other variables are numeric(positive and negative) Although accuracy is same in both cases but I assume that F1 should also be same in both cases. So why it is giving different results. I know that scikit-learn can handle encoding but why F1 is different in both cases?
xtrain,xtest,ytrain,ytest=train_test_split(X,encoded_Y,test_size=0.3,random_state=100,shuffle=True)
clf_xgb = xgb.XGBClassifier(nthread=1,random_state=100)
clf_xgb.fit(xtrain, ytrain)
xgb_pred = clf_xgb.predict(xtest)
xgb_pred_prb=clf_xgb.predict_proba(xtest)[:,1]
print(confusion_matrix(xgb_pred,ytest))
# [984 57]
# [103 1856]
#Find Accuracy of XGBoost
accuracy_xgb = accuracy_score(ytest,xgb_pred)
print("Accuracy: {}".format(accuracy_xgb)
#Find Recall of XGBoost
recall_xgb = recall_score(ytest,xgb_pred)
recall_xgb
#Find Precision of XGBoost
precision_xgb = precision_score(ytest,xgb_pred)
precision_xgb
#Find F1 Score XGB
xgb_f1=f1_score(ytest,xgb_pred)
xgb_f1