Where do Workers and Parameter Servers reside in Distributed TensorFlow?

Viewed 866

In this post, it was mentioned that:

Also, there's no built-in distinction between worker and ps devices -- it's just a convention that variables get assigned to ps devices, and ops are assigned to worker devices.

In this post, it was mentioned that:

TL;DR: TensorFlow doesn't know anything about "parameter servers", but instead it supports running graphs across multiple devices in different processes. Some of these processes have devices whose names start with "/job:ps", and these hold the variables. The workers drive the training process, and when they run the train_op they will cause work to happen on the "/job:ps" devices, which will update the shared variables.

Questions:

  1. Do variables in ps reside on the CPU or GPU? Also, are there any performance gains if "/job:ps" resides on CPU or GPU?
  2. Do the lower level libraries decide where to place a variable or operation?
1 Answers

Do variables in ps reside on the CPU or GPU? Also, are there any performance gains if "/job:ps" resides on CPU or GPU?

You can pin ps job to either on of those (with exceptions, see below), but pinning it to GPU is not practical. ps is really a storage of parameters and ops to update it. A CPU device can have a lot more memory (i.e., main RAM) than a GPU and is fast enough to update the parameters as the gradients are coming in. In most cases, matrix multiplications, convolutions and other expensive ops are done by the workers, hence a placement of a worker on a GPU makes sense. A placement of a ps to a GPU is a waste of resources, unless a ps job is doing something very specific and expensive.

But: Tensorflow does not currently have a GPU kernel for integer variables, so the following code will fail when Tensorflow tries to place the variable i on GPU #0:

with tf.device("/gpu:0"):
  i = tf.Variable(3)

with tf.Session() as sess:
  sess.run(i.initializer)   # Fails!

with the following message:

Could not satisfy explicit device specification '/device:GPU:0' 
because no supported kernel for GPU devices is available.

This is the case when there's no choice of device for a parameter, and thus for a parameter server: only CPU.

Do the lower level libraries decide where to place a variable or operation?

If I get this question right, node placement rules are pretty simple:

  • If a node was already placed on a device in a previous run of the graph, it is left on that device.
  • Else, if the user pinned a node to a device via tf.device, the placer places it on that device.
  • Else, it defaults to GPU #0, or the CPU if there is no GPU.

Tensorflow whitepaper describes also a dynamic placer, which is more sophisticated, but it's not part of the open source version of tensorflow right now.

Related