I have a model that produces parameters for a beta distribution and would like to optimize according to some loss involving the pdf and the cdf of this distribution. However, I recently found that tensorflow's autograd routines are not producing the full gradients as I expected.
Before getting into the details of the example, the TL;DR question is: how can I compute gradients of tensorflow probability distributions with respect to their hyperparameters, e.g., alpha and beta in a beta distribution?
By way of example consider the following simple model (note this was thrown together to illustrate the code, the model itself doesn't make much sense). The model takes 3-dimensional data and produces 2-dimensional outputs. These outputs are shifted and viewed as alpha and beta in a beta distribution. For illustrative purposes I generate some data and aim to maximize the average probability. The following code produces some gradients:
import tensorflow as tf
import tensorflow_probability as tfp
# a dummy model that will take in 3-dimensional data and output 2-dimensional data
model = tf.keras.Sequential()
model.add(tf.keras.layers.Dense(2,activation='relu'))
# data to feed into the model
input_data = tensorflow.random.normal((32,3))
# data to evaluate the distributions on
dist_data = tensorflow.random.uniform((32,1))/2.0+0.25
with tensorflow.GradientTape() as tape:
# generate beta distributions from output of the network
output = model(input_data)
alpha = output[:,0,None]+1
beta = output[:,1,None]+1
dists = tfp.distributions.Beta(alpha,beta)
# compute average probabilty
dist_output = tensorflow.reduce_mean(dists.prob(dist_data))
grads = tape.gradient(dist_output,model.trainable_variables)
print(grads)
Output
[<tf.Tensor: shape=(3, 2), dtype=float32, numpy=
array([[-0.02362663, 0.05741577],
[ 0.02077126, 0.03589955],
[ 0.04922515, -0.0180204 ]], dtype=float32)>, <tf.Tensor: shape=(2,), dtype=float32, numpy=array([0.01512624, 0.08244684], dtype=float32)>]
Now if I instead target the CDF rather than the PDF I get None for gradients:
with tensorflow.GradientTape() as tape:
# generate beta distributions from output of the network
output = model(input_data)
alpha = output[:,0,None]+1
beta = output[:,1,None]+1
dists = tfp.distributions.Beta(alpha,beta)
# compute average cdf
dist_output = tensorflow.reduce_mean(dists.cdf(dist_data))
grads = tape.gradient(dist_output,model.trainable_variables)
print(grads)
Output
[None, None]
I went to the underlying tensorflow_probability code and it appears that tensorflow_probability uses betainc from tensorflow.python.ops.gen_math_ops and that the gradients are not evaluated there. This can be seen by replacing the CDF above with a direct all to the incomplete beta function (grads still return None).
with tensorflow.GradientTape() as tape:
# generate beta distributions from output of the network
output = model(input_data)
alpha = output[:,0,None]+1
beta = output[:,1,None]+1
# compute average cdf
dist_output = tensorflow.reduce_mean(tensorflow.math.betainc(alpha,beta,dist_data))
grads = tape.gradient(dist_output,model.trainable_variables)
print(grads)
Output
[None, None]
Similarly, the PDF code above calls lbeta which also doesn't seem to be considered by the gradient, though in this case 0 is returned instead of None
with tensorflow.GradientTape() as tape:
# generate beta distributions from output of the network
output = model(input_data)
alpha = output[:,0,None]+1
beta = output[:,1,None]+1
# compute average probabilty
norm_output = tensorflow.reduce_mean(tensorflow.math.lbeta(alpha,beta))
norm_grads = tape.gradient(norm_output,model.trainable_variables)
print(norm_grads)
Output
[<tf.Tensor: shape=(3, 2), dtype=float32, numpy=
array([[0., 0.],
[0., 0.],
[0., 0.]], dtype=float32)>, <tf.Tensor: shape=(2,), dtype=float32, numpy=array([0., 0.], dtype=float32)>]
The net result of this is that if we have a complicated objective function involving both the PDF and the CDF, then we get non-zero gradients, but they are not correct as certain terms are not differentiated.
Question: What are my options for computing these gradients in the tensorflow framework?