I am trying to compute an f1 score using true positive and false positive detections (and also eventually false negative detections).
However, I am unsure if I am implementing determining true positives and false positives correctly for future use.
I have amended a part of a code from the following repo that determines the mean average precision:
The only difference is I want to save the true positive detections for later so I can compute the f1 score and also know exact detections and coordinates.
Therefore, I append the detection into a false positive or true positive list.
In this case, the detection with the best "iou" (greater than 0.5) is saved as a true positive, and hopefully everything else is saved as a false positive.
Does this seem correct?:
global_fp = []
global_tp = []
# detection[0] indicates image #
# ground_truth_image: the gt bbox's that are in same image as detection
ground_truth_image = [bbox for bbox in ground_truths if bbox[0] == detection[0]]
# num_gt_boxes: number of ground truth boxes in given image
num_gt_boxes = len(ground_truth_image)
best_iou = 0
best_gt_index = 0
for index, gt in enumerate(ground_truth_image):
iou = torchvision.ops.box_iou(torch.tensor(detection[3:]).unsqueeze(0),
torch.tensor(gt[3:]).unsqueeze(0))
if iou > best_iou:
best_iou = iou
best_gt_index = index
if best_iou > 0.5:
# check if gt_bbox with best_iou was already covered by previous detection with higher confidence score
# amount_bboxes[detection[0]][best_gt_index] == 0 if not discovered yet, 1 otherwise
if amount_bboxes[detection[0]][best_gt_index] == 0:
true_Positives[detection_index] = 1
amount_bboxes[detection[0]][best_gt_index] == 1
global_tp.append(detection)
else:
false_Positives[detection_index] = 1
global_fp.append(detection)
else:
false_Positives[detection_index] = 1
global_fp.append(detection)
As an FYI, an example detection looks like:
[9, 1, 0.622536838054657, 1.0155895948410034, 342.9746398925781, 18.616594314575195, 373.5692443847656]
where the first dimension is the image id, the second is the class label, the third is a probability score and the rest are bounding box coordinates.
and a ground truth (gt) detection looks like:
[310, 1, 1, 271.0, 322.0, 303.0, 354.0]
where everything is the same (the score is given as 1 for all ground truth, however this is never used and is just there to fill in the dimension).