Backgroud:
I use PytorchLightning to train with FasterRCNN and MaskRCNN for object detection with 2 different classes (multi-class classification).
Task:
Now i want to implement metrics e.g Precision, Recall etc. from TorchMetrics.
As predicted output i have for each image a dictionary:
- 'boxes': Tensor with shape (Num_boxes, 4 for box coordinations)
- 'labels': Tensor with shape (Num_boxes,) label could be 1 or 2
- 'scores': the confidence score for each box, Tensor with shape (Num_boxes,)
Problem:
However, the desired input shape for TorchMetric is totally different than what i have as prediction result, the questions are:
- I don't understand what the preds shape is for multi-class (N,) or for multi-class with logits (N,C) and
- how can i convert my output into this shape?
I have a hunch that they may want the box prediction result, i should first filter my output by score threshold and by IoU threshold, like in this function from mmdetection, but after that i can get fp, tp etc. instead of a tensor of shape(N,) or (N,C).
I hope I have stated my problem in an understandable way, any help would be really appreciated!