How to correctly implement dropout for convolution in TensorFlow

Viewed 2356

According to the original paper on Dropout said regularisation method can be applied to convolution layers often improving their performance. TensorFlow function tf.nn.dropout supports that by having a noise_shape parameter to allow the user to choose which parts of the tensors will drop out independently. However, neither the paper nor the documentation give a clear explanation of which dimensions should be kept independently, and the TensorFlow explanation of how noise_shape works is rather unclear.

only dimensions with noise_shape[i] == shape(x)[i] will make independent decisions.

I would assume that for a typical CNN layer output of the shape [batch_size, height, width, channels] we don't want individual rows or columns to drop out by themselves, but rather whole channels (which would be equivalent to a node in a fully connected NN) independently of the examples (i.e. different channels could be dropped for different examples in a batch). Am I correct in this assumption?

If so, how would one go about implementing dropout with such specificity using the noise_shape parameter? Would it be:

noise_shape=[batch_size, 1, 1, channels]

or:

noise_shape=[1, height, width, 1]
1 Answers
Related