Tensorflow: OP_REQUIRES failed at save_restore_v2_ops.cc:184 : Not found: Key global_step not found in checkpoint

Viewed 187

I have been trying to load a pretrained model and use it for further training. But the model I have downloaded is not working for me. Even after reading a lot about this issue I am still lost.

The error is:

I0616 16:24:36.320715 139841803228928 mpi.py:337] parameter_count = 16892227 INFO:tensorflow:Resume training from previous checkpoint I0616 16:24:36.326457 139841803228928 mpi.py:343] Resume training from previous checkpoint WARNING:tensorflow:From /usr/local/lib/python2.7/dist-packages/tensorflow/python/training/saver.py:1276: checkpoint_exists (from tensorflow.python.training.checkpoint_management) is deprecated and will be removed in a future version. Instructions for updating: Use standard file APIs to check for files with this prefix. W0616 16:24:36.326805 139841803228928 deprecation.py:323] From /usr/local/lib/python2.7/dist-packages/tensorflow/python/training/saver.py:1276: checkpoint_exists (from tensorflow.python.training.checkpoint_management) is deprecated and will be removed in a future version. Instructions for updating: Use standard file APIs to check for files with this prefix. INFO:tensorflow:Restoring parameters from /home/stereo_magnification/stereo-magnification-master/model/siggraph_model_20180701/model.latest I0616 16:24:36.330070 139841803228928 saver.py:1280] Restoring parameters from /home/stereo_magnification/stereo-magnification-master/model/siggraph_model_20180701/model.latest 2022-06-16 16:24:36.548759: W tensorflow/core/framework/op_kernel.cc:1502] OP_REQUIRES failed at save_restore_v2_ops.cc:184 : Not found: Key global_step not found in checkpoint INFO:tensorflow:Error reported to Coordinator: <class 'tensorflow.python.framework.errors_impl.NotFoundError'>, Restoring from checkpoint failed. This is most likely due to a Variable name or other graph key that is missing from the checkpoint. Please ensure that you have not altered the graph expected based on the checkpoint. Original error:

Key global_step not found in checkpoint [[node save/RestoreV2 (defined at home/stereo_magnification/stereo-magnification-master/stereomag/mpi.py:329) ]]

Please let me know how to proceed with this.

Here is the pretrained model: https://drive.google.com/open?id=1CZGJxRl0GK0js0MbL7cn7tHtdRrtnjOB

Part of the script:

def train(self, train_op, checkpoint_dir, continue_train, summary_freq,
        save_latest_freq, max_steps):
"""Runs the training procedure.

Args:
  train_op: op for training the network
  checkpoint_dir: where to save the checkpoints and summaries
  continue_train: whether to restore training from previous checkpoint
  summary_freq: summary frequency
  save_latest_freq: Frequency of model saving (overwrites old one)
  max_steps: maximum training steps
"""
parameter_count = tf.reduce_sum(
    [tf.reduce_prod(tf.shape(v)) for v in tf.trainable_variables()])
global_step = tf.Variable(0, name='global_step', trainable=False)
incr_global_step = tf.assign(global_step, global_step + 1)
saver = tf.train.Saver(
    [var for var in tf.model_variables()] + [global_step], max_to_keep=10)
sv = tf.train.Supervisor(
    logdir=checkpoint_dir, save_summaries_secs=0, saver=None)

with sv.managed_session() as sess:
  tf.logging.info('Trainable variables: ')
  for var in tf.trainable_variables():
    tf.logging.info(var.name)
  tf.logging.info('parameter_count = %d' % sess.run(parameter_count))
  if continue_train:
    
    checkpoint = tf.train.latest_checkpoint(checkpoint_dir)
    if checkpoint is not None:
      tf.logging.info('Resume training from previous checkpoint')
      tf.reset_default_graph()
      saver.restore(sess, checkpoint)

This is the original project: https://github.com/google/stereo-magnification

0 Answers
Related