How to evaluate object detection in Python?

Viewed 338

I have a list of ground truth and a list of corresponding prediction for object detection for a given picture in the following form:

ground_truth = [[0,6,234,45,362],
                [1,1,156,103,336],
                [1,36,111,198,416],
                [1,91,42,338,500]]

prediction = [[0,6,234,39,128],
              [0,3,244,39,128]
              [1,1,156,102,180],
              [1,36,111,162,305]]

where the individual objects are descripted as (sub)lists of class, min_x, min_y, max_x, max_y values for the bounding boxes for each recognized objects.

Now, I would like to have an evaluation of the correctness of the predictions against the ground truth. I know about intersection over union and F1 computed from it, but only for one class and one object. I wonder how to generalize this over multiple possible objects of multiple possible classes?

Should I iterate over each ground truth objects and iterate over all prediction objects and see which prediction is the closest? But what if a prediction partly overlaps two ground truth objects? It's confusing...

2 Answers

You can use the SequenceMatcher class from the built-in difflib module to compare two arrays. Your nested lists will need to be flattened for this to work, which can be done with nested simple list comprehensions:

from difflib import SequenceMatcher

ground_truth = [[0,6,234,45,362],
                [1,1,156,103,336],
                [1,36,111,198,416],
                [1,91,42,338,500]]

prediction = [[0,6,234,39,128],
              [0,3,244,39,128],
              [1,1,156,102,180],
              [1,36,111,162,305]]

lst1 = [a for b in ground_truth for a in b]
lst2 = [a for b in prediction for a in b]

print(SequenceMatcher(None, lst1, lst2).ratio())

Output:

0.45

As you can see, the similarity between the two lists is 45%.

Have a look at

sklearn.metrics.f1_score(y_true, y_pred, *, labels=None, pos_label=1, average='binary', sample_weight=None, zero_division='warn')

F-1 score is the harmonic mean of precision and recall.

The argument average will help you deal with multi-class/multi-label cases efficiently: Macro/Micro/Samples/Weighted/Binary are used in the context of multiclass/multilabel targets. If None, the scores for each class are returned. Otherwise, this determines the type of averaging performed on the data:

binary: Only report results for the class specified by pos_label. This is applicable only if targets (y_{true,pred}) are binary.

micro: Calculate metrics globally by counting the total true positives, false negatives and false positives.

macro: Calculate metrics for each label, and find their unweighted mean. This does not take label imbalance into account.

weighted: Calculate metrics for each label, and find their average weighted by support (the number of true instances for each label). This alters ‘macro’ to account for label imbalance; it can result in an F-score that is not between precision and recall.

samples: Calculate metrics for each instance, and find their average (only meaningful for multilabel classification where this differs from accuracy_score)

Examples:

from sklearn.metrics import f1_score
y_true = [0, 1, 2, 0, 1, 2]
y_pred = [0, 2, 1, 0, 0, 1]
f1_score(y_true, y_pred, average='macro')
0.26...
f1_score(y_true, y_pred, average='micro')
0.33...
f1_score(y_true, y_pred, average='weighted')
0.26...
f1_score(y_true, y_pred, average=None)
array([0.8, 0. , 0. ])
y_true = [0, 0, 0, 0, 0, 0]
y_pred = [0, 0, 0, 0, 0, 0]
f1_score(y_true, y_pred, zero_division=1)
1.0...

For a detailed academic paper: A Survey on Performance Metrics for Object-Detection Algorithms

Related