I am working on a multilabel project with 2 labels (the values of the labels can only be 1 or 0) both imbalanced (small amount of 1s compared to 0s). I used Binary Relevance with GaussianNB. However, the f1 micro score was very low (around 0.25 to 0.35) so I tried it instead with only one of the labels. The f1 micro score then went up (around 0.75 to 0.9). I tried to do the same thing manually, and created 2 of the same models just training and testing different labels. They both had a high f1 micro score (around 0.75 to 0.9) which made no sense. I removed one of the labels and added one which always be 1, and the Binary Relevance model had a high f1 micro score (around 0.9) as well as the models that I manually built. Shouldn't the Binary Relevance model have around the same f1 score as the average score of my manually built models?