How to deal with different input sizes in CNN models

Viewed 1602

To give a bit of a context: I'm fairly new to machine learning, I've read and seen some educational videos on how CNN works.

I've tried two models so far, a random person's CNN model and the Google's Inception v3 model. I could understand that random's person CNN model and what's happening in there. What I don't understand is how to make it work with different output sizes that are not just a different scale or rotation. Let me just explain what I'm doing:

I basically want to be able to classify a picture (containing a logo) as a brand. For example, you give me a picture that contains the Starbucks logo and our model will tell you it's Starbucks. There is going to be only one logo in every picture (for my case). First try was with the inception model: tried with 20,000 iterations with 2,000 Starbucks receipt pictures, 2,000 Walmart receipt pictures and 2,000 random pictures that were not related to Starbucks or Walmart so I could also classify the picture as 'Neither'. Got 88% accuracy, not good enough and the cross entropy doesn't drop to lower than 0.4 then I tried cropping the logo from those picture and tried again. this time, on cropped pictures it would work like a charm but on bigger pictures containing the starbucks logo, or walmart for that matter, it would fail miserably.

Same thing with the DeepLogo's way: https://github.com/satojkovic/DeepLogo

It works well with the 32 x 32 picture but once I change the input size, it fails.

How can I overcome this?

EDIT: I'm using this for retraining on top of the Inception model: https://github.com/tensorflow/tensorflow/tree/master/tensorflow/examples/image_retraining

1 Answers
Related