Best practice in storing large number of images for training using Google ML engine and Google Storage

Viewed 592

I'm training an SSD model in TensorFlow using Google ML engine and Google Storage. In TF's object detection example, they put all images into a single big TFRecord file. However, in this scheme, if one wants to assemble different training set by choosing subset of all images, a given image will be stored multiple times, once for each training set the image belongs to.

The alternative is to store each image as an individual file, and use a flat list of URLs to indicate the membership of an image in various data sets. However, based on my experience, Google Storage isn't optimized for reading large number of small files, which results in low training throughput.

I would like to see if there're other ways to avoid saving each image multiple times while achieve good throughput.

2 Answers

Small files on GCS do hurt throughput.

A few ideas:

  1. Build your input pipeline with many reading threads to keep the pipe full. (Link to newer API)
  2. Copy the files to local disk at startup.
  3. Use constructs in your TF graph to filter out files.

No. 1 should get you pretty far.

Since having a large number of files reduces training throughput, what i would do is:

  • Put the images in a large tfrecord. The record would be setup that one of the fields would a subset key.
  • Using the new Dataset API I would load only the required dataset using an appropriate parse function.

Assuming that the images are shuffled appropriately, the subset you are choosing from is big enough and considerable reading threads are used, as the pipeline shouldn't run out of data.

Another approach would be to divide the tfrecords into smaller subsets but not files for each image. Either way you go, you will have some issues you need to tackle, this is a case to choose which option has less possible problems.

Related