Tensorflow: How do you monitor GPU performance during model training in real-time?

Viewed 21290

I am new to Ubuntu and GPUs and have recently been using a new PC with Ubuntu 16.04 and 4 NVIDIA 1080ti GPUs in our lab. The machine also has an i7 16 core processor.

I have some basic questions:

  1. Tensorflow is installed for GPU. I presume then, that it automatically prioritises GPU usage? If so, does it use all 4 together or does it use 1 and then recruit another if needed?

  2. Can I monitor in real-time, the GPU use/activity during training of a model?

I fully understand this is basic hardware stuff but clear definitive answers to these specific questions would be great.

EDIT:

Based on this output - it this really saying that nearly all the memory on each one of my GPUs is being used?

enter image description here

7 Answers

I would suggest nvtop, it shows real-time status and easier to watch than nvidia-smi. It also shows in a graph.

$ sudo apt install nvtop
$ nvtop

enter image description here

Try this command:

nvidia-smi --query-gpu=utilization.gpu --format=csv --loop=1

Here is a demo:

enter image description here

All the above commands use watch, it's much more efficient to keep the context alive by using the builin looper: nvidia-smi -l 1.

If you want to see something like htop and nvidia-smi at the same time, you can try glances (pip install glances).

You should use nvidia-smi. Just keep in mind that depending on your workload you might not see any change in the load if the task completes between 2 sampling events.

Also keep in mind that the maximum sampling interval is 1/6 second as per: http://manpages.org/nvidia-smi

Utilization rates report how busy each GPU is over time, and can be used to determine how much an application is using the GPUs in the system. Note: During driver initialization when ECC is enabled one can see high GPU and Memory Utilization readings. This is caused by ECC Memory Scrubbing mechanism that is performed during driver initialization.

GPU Percent of time over the past sample period during which one or more kernels was executing on the GPU. The sample period may be between 1 second and 1/6 second depending on the product.

Memory Percent of time over the past sample period during which global (device) memory was being read or written. The sample period may be between 1 second and 1/6 second depending on the product.

Related