I have built a model with multiple heads, some doing regression and some classification. I then sum up all the losses in a weighted manner to backpropagate.
For the classification heads, I use the one-hot encoding approach, and code takes argmax on the model outputs to get the class. I have found that the Categorical Crossentropy loss (or BCE) doesn't work and the outputs are mostly homogeneous and uniform, meaning the model is not learning. However, a simple change to MSE loss gives good results. Can you tell me why this might be happening?
I have tried a number of last-layer activation combinations, MSE gives good results with no last layer activation.
An example of one-hot vectors I am trying to learn:
|-------------|----------| |-------------|-----------------|
| X | y | | X | One-Hot-Y |
|-------------|----------| |-------------|-----------------|
| DP1 | A | | DP1 | [1, 0, 0, 0, 0] |
|-------------|----------| |-------------|-----------------|
| DP2 | C | | DP2 | [0, 0, 1, 0, 0] |
|-------------|----------| |-------------|-----------------|
| DP3 | E | | DP3 | [0, 0, 0, 0, 1] |
|-------------|----------| |-------------|-----------------|
| DP4 | A | | DP4 | [1, 0, 0, 0, 0] |
|-------------|----------| |-------------|-----------------|
| DP5 | D | | DP5 | [0, 0, 0, 1, 0] |
|-------------|----------| |-------------|-----------------|