Why set num_batch_threads to number CPU cores with GPU setup

Viewed 185

With Tensorflow Serving, the recommendation for batching parameters for GPU inference is to set num_batch_threadsto the # of CPU cores https://github.com/tensorflow/serving/blob/master/tensorflow_serving/batching/README.md#gpu-one-approach. However, experimentally, I've seen that setting num_batch_threads=1 has produced lower latencies at lower rps (i.e. 200). I'm using a 12 cores of CPU and 1 V100 GPU. I wanted to understand why this recommendation was made and how to choose a good value for this parameter.

0 Answers
Related