I am developing a multilabel classification model using keras.
If I use sigmoid as the last activation function with binary cross-entropy loss, I get 98% of accuracy in my first epoch, but its actually not learning anything. It is because I have 800 categories and have an average of 2-3 True positives.
So I changed to using sigmoid with categorical cross-entropy loss and it seems to work better but I think it is not recommended?
I would like some guidance on what activation and loss I should use in this case. Thanks!