I'm trying to evaluate the performance of an object detection model using torchmetrics mean average precision. However, I'm getting this odd result:
When I evaluate the metric for image A, I get 'map': 0.7891, 'map_50': 1, 'map_75': 1
When I evaluate the metric for image B, I get 'map': 0.7611, 'map_50': 1, 'map_75': 1
When I run the metric for a list containing information about both image A and image B, I expect the performance to be somewhere between 0.7891 and 0.7611. Instead, I get 'map':0.7197, 'map_50': 0.9293,'map_75': 0.9293.
Is my reasoning about how mean average precision for multiple images works flawed, or is something going wrong here?
Here are the specific values I'm using:
import torch
from torchmetrics.detection.mean_ap import MeanAveragePrecision
include_A = True
include_B = True
truth_eval = []
pred_eval = []
if include_A:
truth_eval.append({'boxes': torch.tensor([[562., 180., 694., 227.],
[322., 189., 467., 251.],
[770., 167., 954., 240.]], dtype=torch.float64),
'labels': torch.tensor([1., 1., 1.], dtype=torch.float64)})
pred_eval.append({'boxes': torch.tensor([[ 769.4067, 163.4746, 949.1498, 240.5295],
[ 574.9888, 178.3064, 697.7145, 226.6653],
[ 333.8886, 188.9244, 471.6525, 250.9485],
[1037.9913, 20.9808, 1110.0175, 118.0723]]),
'scores': torch.tensor([1., 1., 1., 1.]),
'labels': torch.tensor([1., 1., 1., 1.])})
if include_B:
truth_eval.append({'boxes': torch.tensor([[554., 180., 688., 228.],
[309., 190., 457., 253.],
[810., 166., 999., 239.]], dtype=torch.float64),
'labels': torch.tensor([1., 1., 1.], dtype=torch.float64)})
pred_eval.append({'boxes': torch.tensor([[ 567.4130, 177.8370, 691.1442, 227.1231],
[ 810.1462, 163.8314, 994.3475, 241.7314],
[ 321.8360, 190.0476, 464.3474, 251.8763],
[1046.0631, 18.7516, 1111.9734, 104.2167],
[ 540.7168, 76.8780, 564.6191, 133.8904]]),
'scores': torch.tensor([1., 1., 1., 1., 1.]),
'labels': torch.tensor([1., 1., 1., 1., 1.])})
metric = MeanAveragePrecision()
metric.update(pred_eval , truth_eval )
metric.compute()