Tensorflow Object Dectection API : how to create tfrecords with images not containing any labels (hard negatives)?

Viewed 1400

Hello.

I am currently using tensorflow object detection API(with Faster Rcnn)on my own dataset, for some of my labels, i've identified objects that are very likely to be detected as false positives, and i know that the API uses hard exemples mining, so what i'am trying to do it is to introduce images that contain these hard objects into the training , so that the miner can take them as hard negatives.

following this conversation on github https://github.com/tensorflow/models/issues/2544 I was told that it's possible

You can have purely negative images and faster_rcnn models will sample from anchors from them.

So my question is : how do i create tfrecords with some images not having any bounding boxes ? What do i put in the associated .xml files?

3 Answers

In your tfrecords generation script make sure you add the metadata of the hard negative image to tf record like this -

tf_example = tf.train.Example(features=tf.train.Features(feature={
            'image/height': dataset_util.int64_feature(height),
            'image/width': dataset_util.int64_feature(width),
            'image/filename': dataset_util.bytes_feature(filename),
            'image/source_id': dataset_util.bytes_feature(filename),
            'image/encoded': dataset_util.bytes_feature(encoded_jpg),
            'image/format': dataset_util.bytes_feature(image_format)
            }))

For an image with objects you must add the bounding boxes and label information as well -

tf_example = tf.train.Example(features=tf.train.Features(feature={
            'image/height': dataset_util.int64_feature(height),
            'image/width': dataset_util.int64_feature(width),
            'image/filename': dataset_util.bytes_feature(filename),
            'image/source_id': dataset_util.bytes_feature(filename),
            'image/encoded': dataset_util.bytes_feature(encoded_jpg),
            'image/format': dataset_util.bytes_feature(image_format),
            'image/object/bbox/xmin': dataset_util.float_list_feature(xmins),
            'image/object/bbox/xmax': dataset_util.float_list_feature(xmaxs),
            'image/object/bbox/ymin': dataset_util.float_list_feature(ymins),
            'image/object/bbox/ymax': dataset_util.float_list_feature(ymaxs),
            'image/object/class/text': dataset_util.bytes_list_feature(classes_text),
            'image/object/class/label': dataset_util.int64_list_feature(classes),
        }))

Also make sure you set min_negatives_per_image to a positive number in your pipeline.config file otherwise it won't train with negative images

I adapted my dataset to add a dummy annotation for images that doesn't had any real annotations and changed the tfrecord producer code to:

def create_tf_example(group, path, label_map):
    with tf.gfile.GFile(os.path.join(path, '{}'.format(group.filename)), 'rb') as fid:
        encoded_jpg = fid.read()
    encoded_jpg_io = io.BytesIO(encoded_jpg)
    image = Image.open(encoded_jpg_io)
    width, height = image.size

    filename = group.filename.encode('utf8')
    image_format = b'jpg'

    xmins = []
    xmaxs = []
    ymins = []
    ymaxs = []
    classes_text = []
    classes = []

    for index, row in group.object.iterrows():
        if not pd.isnull(row.xmin):
            if not row.xmin == -1:
                xmins.append(row['xmin'] / width)
                xmaxs.append(row['xmax'] / width)
                ymins.append(row['ymin'] / height)
                ymaxs.append(row['ymax'] / height)
                classes_text.append(row['class'].encode('utf8'))
                classes.append(label_map[row['class']])

    tf_example = tf.train.Example(features=tf.train.Features(feature={
        'image/height': dataset_util.int64_feature(height),
        'image/width': dataset_util.int64_feature(width),
        'image/filename': dataset_util.bytes_feature(filename),
        'image/source_id': dataset_util.bytes_feature(filename),
        'image/encoded': dataset_util.bytes_feature(encoded_jpg),
        'image/format': dataset_util.bytes_feature(image_format),
        'image/object/bbox/xmin': dataset_util.float_list_feature(xmins),
        'image/object/bbox/xmax': dataset_util.float_list_feature(xmaxs),
        'image/object/bbox/ymin': dataset_util.float_list_feature(ymins),
        'image/object/bbox/ymax': dataset_util.float_list_feature(ymaxs),
        'image/object/class/text': dataset_util.bytes_list_feature(classes_text),
        'image/object/class/label': dataset_util.int64_list_feature(classes),
    }))
    return tf_example

So, when an dummy annotation 'xmin == -1' appears, it creates a tfrecord with empty list of bounding boxes (class, xmin, xmax, ymin, ymax).

Besides the messy train loss behavior, my model successfully learned the negative samples pattern, thus decreasing to zero the false positives I got, on my scenario.

enter image description here

You don't have to do anything specific for that. Just leave the list of associated bounding boxes empty, whatever the source format is. I did so for my experiments, did not get any gain though.

Related