Fast R-CNN: do we extract region proposals from a single feature map?

Viewed 711

I'm trying to understand the evolution in object detection algorithms. As reference I'm using Andrej Karpathy slides of lecture 8 here the slides.

I think I understood why we go from R-CNN to Fast R-CNN: instead of running each (warped) proposed region into a Convnet we run the original image into a Convnet and then we extract the region proposals from the feature map (or do we? I also wonder how initial RoI are mapped to feature maps):

Here lie my doubts: a single feature map? How do we get from many feature maps to a single one? Or is it just a definition misunderstanding, where with feature map we are considering a tensor with depth equal to the number of 2d feature maps exiting from the convolutional layers?

Context: Usually I have many feature maps as output of a convolutional network, which then I transform into a fully connected layer to do classification. To me a feature map is the output of a convolution (+ relu + pooling) operation, so I have one feature map for each convolutional filter (consider for my understanding a single channel image so that we can refer just to 2d convolution).

Thank you for your help

0 Answers
Related