I have been playing around with custom loss functions for a while with some success, but I'm struggling with a new loss function, and I wonder if it might be due to the loss result tensor's shape.
My y_true and y_pred tensors have shape == (100, 216, 563). Due to the nature of the data and the calculations I'm performing in my loss function, it makes perfect sense to output a loss tensor of shape == (100, 563) because the second dimension gets reduced away with a reduce_prod() operation.
However, if I use this loss function alone, the loss value steadily increases instead of decreasing... I've not seen this before. If it was all over the place I'd think it was just a bad idea for a loss function or my maths was wrong somewhere, but as far as I can tell the maths is right.
Will this weird shape with a missing middle dimension throw off the gradient calculations? I've tried already using keepdims=True in my reduce_foo() methods, but this makes no difference to the increasing loss value (and the results still have a different shape, shape == (100, 1, 563)
Looking through tensorflow docs, I can find examples of both a loss with matching shape to y_pred and y_true, and another loss with a single scalar value. Are there any specific rules stated anywhere as to what shape the output loss should be or can anyone give me insights that might help me understand why the loss should be a specific shape (if that is even my problem)?