I already created an Ensemble for classification (average -or so- of predictions per images), or for Semantic Segmentation (average -or so- of predictions per pixels), but I don't really know how to proceed for Object Detection.. My guess would be to extract all the region proposals of all my networks, then to run my classifiers on the X best of them, and finally to average the predictions for all the bounding boxes. But how should I do that with architectures following the Object Detection API?
I guess the regions proposals can be extracted using extract_proposal_features, and then reinserted to the model, but the only way I see to do that would be to create a complete new model with its own predict method etc, dealing will all the models of my Ensemble. Am I missing an other obvious / simpler method?