How does size of training image, training network and inference network affect the accuracy of YOLOv4 Darknet model?

Viewed 582

Our use case is to train a YOLOv4 network to detect an object as small as a wedding band on top of a table. So the object is about 250px by 250px in a 4096px x 2160px.

  1. Training network size: According to the README.MD, should our training net to be as big as it can fit within the GPU memory? In our case, it's 1056 x 1056 for a Titan T4 with 16GB GPU RAM

  2. Training images size: For our training images of the object, they are around 1000 px by 1000 px, from the Darknet documentation I see that it will resize when random=1, so I assume we are good with the relatively high resolution training images?

  3. Detection network size: During the detection, in the same README.MD, the first point of item 2 suggested increasing the network-resolution for detection. So if we were to use 1056x1056 during training, should we use 1280 x 1280 (32 * 40) or larger net.width and net.height?

Thanks!

0 Answers
Related