I want to train a convolutional neural network (CNN) on a multi-class dataset. There's a class imbalance, so I want to upsample the minority classes.
The dataset looks as follows:
[Text 1, ["obscene", "insult"]],
[Text 2, ["identity_attack", "insult"]]
I tried to use the resample function from sklearn.utils. The problem is that if I upsample the obscene class by creating a duplicate of Text 1, I automatically upsample the insult class too.
Is there a Python function that solves my problem?