I was trying to understand how K.layers.Dropout might be implemented, since in the literature it's always referred as a random independent sampling of 0/1 masks for each element.
Given that the literature it's pretty clear to me, I switched to coding it, and I stumbled upon an issue: since TF uses Graphs, we don't know the sample size, in particular:
def class CustomLayer(K.keras.Layer)
def call(inputs):
tf.print(inputs.shape)
will indeed print (supposing the eager evaluation is turned off) None as first dimension
Having said that, how is TF able to sample an independent mask for each sample in each minibatch?
At the moment my best guess is that they are using something like tf.vectorized_map to get the performance they are getting with a random mask for each element in the minibatch