Sagemaker Instance not utilising GPU during training

Viewed 995

I'm training a Seq2Seq model on Tensorflow on a ml.p3.2xlarge instance. When I tried running the code on google colab, the time per epoch was around 40 mins. However on the instance it's around 5 hours!

This is my training code

def train_model(train_translator, dataset, path, num=8):

  with tf.device("/GPU:0"):
    cp_callback = tf.keras.callbacks.ModelCheckpoint(filepath=path,
                                                save_weights_only=True,
                                                 verbose=1)
    batch_loss = BatchLogs('batch_loss')
    train_translator.fit(dataset, epochs=num,callbacks=[batch_loss,cp_callback])  

  return train_translator

I have also tried without the tf.device command and I still get the same timing. Am I doing something wrong?

3 Answers

I had to force GPU use with the help of

with tf.device('/device:GPU:0')

If you're using SageMaker Notebook instance. Open a terminal and run nvidia-smi to see the GPU utilization rate. If you it's 0% then you're not using the right device. If it's more than 0% but very far from 100%, then you have a non GPU bottleneck to handle.
If you're using SageMaker training, then check the GPU usage via Cloudwatch metrics for the job.

If you are using a Sagemaker Notebook on a GPU instance with instance_type='local' in the Tensorflow estimator, it apparently defaults to CPU...

I solved this by setting: instance_type='local_gpu' instead.

Related