Why is MSE loss working better for multi-class classification problem than categorical crossentropy?

Viewed 169

I have built a model with multiple heads, some doing regression and some classification. I then sum up all the losses in a weighted manner to backpropagate.

For the classification heads, I use the one-hot encoding approach, and code takes argmax on the model outputs to get the class. I have found that the Categorical Crossentropy loss (or BCE) doesn't work and the outputs are mostly homogeneous and uniform, meaning the model is not learning. However, a simple change to MSE loss gives good results. Can you tell me why this might be happening?

I have tried a number of last-layer activation combinations, MSE gives good results with no last layer activation.

An example of one-hot vectors I am trying to learn:

|-------------|----------|              |-------------|-----------------|
|      X      |     y    |              |      X      |    One-Hot-Y    |
|-------------|----------|              |-------------|-----------------|
|     DP1     |   A      |              |     DP1     | [1, 0, 0, 0, 0] |
|-------------|----------|              |-------------|-----------------|
|     DP2     |   C      |              |     DP2     | [0, 0, 1, 0, 0] |
|-------------|----------|              |-------------|-----------------|
|     DP3     |   E      |              |     DP3     | [0, 0, 0, 0, 1] |
|-------------|----------|              |-------------|-----------------|
|     DP4     |   A      |              |     DP4     | [1, 0, 0, 0, 0] |
|-------------|----------|              |-------------|-----------------|
|     DP5     |   D      |              |     DP5     | [0, 0, 0, 1, 0] |
|-------------|----------|              |-------------|-----------------|
0 Answers
Related