Weight parameter for unbalanced data set in logistic regression

Viewed 1867

I am confused about how to choose a correct weight parameter for unbalanced data set. My data are binary variables with only around 4% of the data are '1' and 96% are '0'. I wanted to use logistic regression specifying a weight.

In this link: https://stats.stackexchange.com/questions/164693/adding-weights-to-logistic-regression-for-imbalanced-data The person seems to say that if we want to use 10% of 0's and 100% of 1's, the weight in the glm() function from R, should have value 10 for observations with y=0 and 1 for observations with y=1. I don't understand how these number are choose, as for me the weight should be increase for the sample in the minority class instead (oversampling methods, while the suggest method seems to be downsampling, which I don’t understand how can this be done with a weight parameter).

I am using glm() function, and would like to take into account all the observations, but counting more for the 1's observations (let's say without loss of generality 10 times more).

Thank you very much for your help!

0 Answers
Related