Semantic alignment loss for fine tuning of object detection model on virtual paintings (style transfer) images

Viewed 25

I'm working on an image captioning project for artworks, in particular I'm trying to implement the strategy of the paper Data-efficient image captioning of fine art paintings via virtual-real semantic alignment training.

The step on which I'm currently working is training a object detection model on a dataset of virtual paintings obtained using a style transfer.

In this notebook you can find what I've tried so far.

All datasets and external code required are downloadable through code in the notebook. Unfortunately Google impose some limits on downloads using gdown, maybe in such case it's possible to manually download and upload the files, here the drive links:

To save time in this attempt:

  • no augmentations are performed on the data because I disabled them in the CustomMapper class
  • I start from a ResNet50 model (instead of ResNet101)
  • I train the model for just one epoch

Once the training will be correctly implemented I will work with a more complete setup (10 epoch, ResizeShortestEdge data augmentation and ResNet101 model).

The model that I obtain through this simplified training is downloadable here.

I computed the object detection metrics using the provided model on the coco2017 virtual paintings validation dataset. The results are worse than the pretrained model imported using detectron2 Model Zoo.

My main doubts are about the semantic alignment loss (explained in the paper), according to the paper it should be the sum of Mean Squared Errors computed between features at different levels, so I used as a backbone an FPN network and I compute the semantic alignment loss in the following way:

align_loss = 0
for virtual_painting_features, features in zip(virtual_painting_features_dict.values(), features_dict.values()):
    align_loss += align_loss_fn(virtual_painting_features, features.detach().clone())

I'm not sure if this is the correct interpretation and if the FPN network is a correct choice for the backbone because no indication about this are in the paper.

Any help or suggestion would be really appreciated!!! Thanks in advance!

0 Answers
Related