We are trying to solve a binary classification problem where during training we have kept the two classes roughly equally proportioned 50-50, but during inferencing we will not be having balanced distribution of the labels. We specifically expect to see one of the classes(the positive class, label 1) as a severe minority as compared to the other(the negative class, label 0).
The below table shows how the precision of the minority class falls with the fall in proportion of the minority class(the positive class, label 1).
| S.No. | Label 1 | Label 0 | Percentage of Label 1 | Precision_1 | Recall_1 |
|---|---|---|---|---|---|
| 1 | 50 | 50 | 50 | 0.9574468085106383 | 0.9 |
| 2 | 50 | 500 | 9 | 0.5625 | 0.9 |
| 3 | 50 | 1000 | 4.7 | 0.39823008849557523 | 0.9 |
| 4 | 50 | 2000 | 2.4 | 0.2647058823529412 | 0.9 |
| 5 | 50 | 5000 | 0.9 | 0.1278409090909091 | 0.9 |
We will be doing inference on batches of such data, where the class which is a minority could be ranging from negligible ~0.01% to a much higher number ~50-60%.
While class balancing(giving more weightage to one of the classes during inference) could help if the percentage of the minority class is fixed, this is not case here since the percentage of the minority class is not fixed.
Has anyone faced the same problem and are there any suggestions as to how we can counter this problem ? Please let us know.