I am working on a binary classification problem with an imbalanced dataset. I have decided to downsample the majority class and I’m wondering what the best approach is when calculating performance metrics on a model that has been trained on a downsampled dataset.
I noticed that the sklearn.metrics.precision_score and sklearn.metrics.recal_score functions have a sample_weight attribute. Is the purpose of this attribute to supply a weight for the downsampled class relative to the ratio in which I downsampled?
For example, if I had 1,000,000 samples for the negative class and I decided to downsample to 100,000, would I set the sample_weight attribute to be equal to 1,000,000 / 100,000 = 10 for the negative class?