how to loop over data feeded with a placeholder in tensorflow only one time using the new Dataset API

Viewed 4277

I am starting to use the new dataset API and one thing that I want to do is not described on the doc (https://www.tensorflow.org/programmers_guide/datasets#training_workflows)

My data fit in memory so I want load it in tensorflow to make the training efficient and for this I see for now 2 way to do it:

one is loading the data in the graph directly like this:

dataset = tf.contrib.data.Dataset.from_tensor_slices((X, Y))
iterator = dataset.make_initializable_iterator()
next_element = iterator.get_next()

# loop on epochs
for _ in range(5):
    # Initialize an iterator over the training dataset.
    sess.run(iterator.initializer)
    # loop over all the batch
    for _ in range(1000):
        s = time.time()
        try:
            sess.run(next_element)
        except tf.errors.OutOfRangeError:
            print("Finish epoch")

the other one is to load the data in a placeholder so the data is not save in the graph:

features_placeholder = tf.placeholder(features.dtype, features.shape)
labels_placeholder = tf.placeholder(labels.dtype, labels.shape)

dataset = tf.contrib.data.Dataset.from_tensor_slices((features_placeholder, labels_placeholder))
iterator = dataset.make_initializable_iterator()
next_element = iterator.get_next()

# loop on epochs
for _ in range(5):
    # Initialize an iterator over the training dataset.
    sess.run(iterator.initializer, feed_dict={features_placeholder: X, labels_placeholder: Y})
    # loop over all the batch
    for _ in range(1000):
        s = time.time()
        try:
            sess.run(next_element)
        except tf.errors.OutOfRangeError:
            print("Finish epoch")

The second is I think the best to save memory, but I don't want to feed the data at each epoch. It is really a loss of performance for nothing.

Is there a way to initialize the iterator only one time with a placeholder?

something like this:

sess.run(iterator.initializer, feed_dict={features_placeholder: X, labels_placeholder: Y})

# loop on epochs
for _ in range(5):
    # Initialize an iterator over the training dataset.
    sess.run(iterator.initializer)
    # loop over all the batch
    for _ in range(1000):
        s = time.time()
        try:
            sess.run(next_element)
        except tf.errors.OutOfRangeError:
            print("Finish epoch")

That way we can keep the performance of the first solution and saving memory like the second solution.

Note:

one solution is to define the number of epoch with dataset.repeat() method but with it we kind of loose track of where we are in the training.

I want to check after each epoch (one pass over all the data) the evolution of the loss.

2 Answers

First of all, I would recommend quantifying the performance overhead of feeding X and Y each time you initialize the iterator. For primitive types like tf.int32 and tf.float32 it is often possible to feed a value without copying any data, and in that case the overhead will be negligible. Even if a copy is necessary, it will entail a single memcpy(), which can be surprisingly fast. (On the other hand, feeding a tf.string tensor can be more expensive, because it requires multiple small copies to convert between the Python and C++ string representations.)

Assuming that it is a significant overhead, you can make it a one-time cost by storing the input data in a tf.Variable. For example:

placeholder_X = tf.placeholder(X.dtype, X.shape)
var_X = tf.Variable(placeholder_X)
placeholder_Y = tf.placeholder(Y.dtype, Y.shape)
var_Y = tf.Variable(placeholder_Y)

dataset = tf.contrib.data.Dataset.from_tensor_slices((var_X, var_Y))
iterator = dataset.make_initializable_iterator()

# ...

# The contents of `X` and `Y` will be copied once, in this call.
sess.run(tf.global_variables_initializer(), feed_dict={
    placeholder_X: X, placeholder_Y = Y})

for _ in range(5):
  # The iterator will be initialized from the variables with no copy.
  sess.run(iterator.initializer)

  # ...

I don't think you need to initialise at every epoch. You can do it once before the training loop. But you also need to tell the data set to repeat and reshuffle at each iteration when you are defining the dataset:

features_placeholder = tf.placeholder(features.dtype, features.shape)
labels_placeholder = tf.placeholder(labels.dtype, labels.shape)

dataset = tf.data.Dataset.from_tensor_slices((features_placeholder, labels_placeholder)).shuffle(buffer_value,reshuffle_each_iteration=True).repeat().batch(batch_num)
iterator = dataset.make_initializable_iterator()
next_element = iterator.get_next()

#Initialize an iterator over the training dataset.
sess.run(iterator.initializer, feed_dict={features_placeholder: X, labels_placeholder: Y})
    # loop on epochs
for _ in range(5):
    # loop over all the batch
    for _ in range(1000):
        s = time.time()
        sess.run(next_element)

Note that since the repeat is on, you need to compute the exact number of iterations that you need to loop over data once in each epoch and set it in your inner loop.

Related