I am doing traffic sign recognition work using German Traffic Sign Detection Benchmark database. This has 43 classes, with at least 400 images in each class. Images may have up to 3 traffic signs.
When I have randomly selected images for training and validation set I get a huge difference in network's test accuracy. I constructed two data sets: one has 75% training images and 25% validation images; the other has 70% training images and 30% validation images.
I am using GoogLeNet with identical hyper-parameters for training, including 30 epochs.
After training, I test with a different set of images that are designed for testing. With the first data set, I get almost 10% lower accuracy than with the second one. Could some one explain this?
Could it be that it randomly selected "easier" images for training and that is why I am getting lower results?
P.S. for both of the data sets I am using the same images, just dividing it differently by percentages.
Link to data set: http://benchmark.ini.rub.de/?section=gtsrb&subsection=dataset