Object detection loss not decreasing, AP increasing

Viewed 1709

I'm trying to train an object detection model to detect and classify 10 classes. My original dataset is pretty sparse and unbalanced, containing a total of 3k tagged images with the following distribution between classes:

Class 1: 21

Class 2: 22

Class 3: 9

Class 4: 192

Class 5: 2240

Class 6: 319

Class 7: 56

Class 8: 190

Class 9: 44

Class 10: 167

Because of this sparsity, I have performed augmentation on all images, namely adding noise, blur, contrast, brightness, and horizontal flipping. I also augmented the flipped images further with noise, contrast, and brightness. The resulting dataset consists of 37k tagged images with the following distribution:

Class 1: 4235

Class 2: 5365

Class 3: 2385

Class 4: 10755

Class 5: 17185

Class 6: 4035

Class 7: 3150

Class 8: 3820

Class 9: 555

Class 10: 1500

The image below shows the different losses for 4 different sessions. The pink graph is the result from the augmented dataset of 37k images, and the other graphs are from previous runs on the original dataset of ~2.5k images. As you can see from the pink graph, the total loss is not decreasing at all from its initial value (as is the case for the blue and red graph from previous runs). The RPN loss is decreasing, but the box classifier loss is increasing, what can be the reason for this?

Losses for different training sessions

I have also included an image of the average precision for each class. The fact that the precision for most classes increases steadily the whole time while the loss is not decreasing seems to me like the model is overfitting? Is it a bad idea to 10x the dataset by augmenting every like I've done? I've also included the config file I'm using below. Any suggestions as to how to improve my training result is appreciated!

Average precision for classes

 model {
  faster_rcnn {
    num_classes: 10
    image_resizer {
     fixed_shape_resizer {
       height: 300
       width: 500 
     }
    }
    feature_extractor {
      type: 'faster_rcnn_inception_resnet_v2'
      first_stage_features_stride: 8
    }
    first_stage_anchor_generator {
      grid_anchor_generator {
        scales: [0.25, 0.5, 0.75, 1.0, 1.5,2,3]
        aspect_ratios: [0.5,1,2,3]
        height: 32
        width: 32
        height_stride: 8
        width_stride: 8
      }
    }
    first_stage_atrous_rate: 1
    first_stage_box_predictor_conv_hyperparams {
      op: CONV
      regularizer {
        l2_regularizer {
          weight: 0.0
        }
      }
      initializer {
        truncated_normal_initializer {
          stddev: 0.01
        }
      }
    }
    first_stage_nms_score_threshold: 0.0
    first_stage_nms_iou_threshold: 0.5
    first_stage_max_proposals: 300
    first_stage_localization_loss_weight: 2.0
    first_stage_objectness_loss_weight: 1.0
    initial_crop_size: 17
    maxpool_kernel_size: 1
    maxpool_stride: 1
    second_stage_box_predictor {
      mask_rcnn_box_predictor {
        use_dropout: True
        dropout_keep_probability: 0.6
        fc_hyperparams {
          op: FC
          regularizer {
            l2_regularizer {
              weight: 0.0
            }
          }
          initializer {
            variance_scaling_initializer {
              factor: 1.0
              uniform: true
              mode: FAN_AVG
            }
          }
        }
      }
    }
    second_stage_post_processing {
      batch_non_max_suppression {
        score_threshold: 0.3
        iou_threshold: 0.5
    #    soft_nms_sigma: 0.5
    #   use_class_agnostic_nms: True
    #    max_classes_per_detection: 1
        max_detections_per_class: 100
        max_total_detections: 100
      }
      score_converter: SOFTMAX
    }
    second_stage_localization_loss_weight: 2.0
    second_stage_classification_loss_weight: 1.0
  }
}
train_config: {
  batch_size: 1
  use_multiclass_scores : False  
  optimizer {
    #momentum_optimizer: {
    adam_optimizer: {
      learning_rate: {
        manual_step_learning_rate {
          initial_learning_rate: 0.0001
          schedule {
            step: 150000
            learning_rate: .00001
          }
          schedule {
            step: 250000
            learning_rate: .000001
          }
        }
      }
      #momentum_optimizer_value: 0.9
    }
    use_moving_average: false
  }
  gradient_clipping_by_norm: 10.0
  from_detection_checkpoint: false

  data_augmentation_options {
    random_horizontal_flip {}
  }

  data_augmentation_options {
    random_crop_image {
      min_object_covered : 1.0
      min_aspect_ratio: 1
      max_aspect_ratio: 1
      min_area: 0.5
      max_area: 1
       random_coef: 0.5
     }
  }
}
1 Answers

Unfortunately with deep learning, it can often times be difficult to distinguish what exact parameter is causing you an issue. From looking at your question, it seems that even with the data augmentation, which is a great thing to do, is that the number of images that you have per class varies very heavily.

For example after you augment your data, you end up with these to classes with images in them

Class 5: 17185

Class 9: 555

Class 5 has 17,185 images, while class 9 only has 555. There is a huge imbalance in the number of images, and its often times preferred to have as close as possible the same amount of images for each class.

When you are training, you will have a validation section, where a pool of images from all classes will be used to test the model at that instance. If you have a bunch of images in one class and not the other, the model will do fairly better validating the images from the larger class, while struggling with the validation of the image from the smaller class because the model didn't have many examples to train with, and or it started to train heavier with the larger class because there are more training examples in there.

Related