Computing gradients of distribution parameters in tensorflow probability

Viewed 248

I have a model that produces parameters for a beta distribution and would like to optimize according to some loss involving the pdf and the cdf of this distribution. However, I recently found that tensorflow's autograd routines are not producing the full gradients as I expected.

Before getting into the details of the example, the TL;DR question is: how can I compute gradients of tensorflow probability distributions with respect to their hyperparameters, e.g., alpha and beta in a beta distribution?

By way of example consider the following simple model (note this was thrown together to illustrate the code, the model itself doesn't make much sense). The model takes 3-dimensional data and produces 2-dimensional outputs. These outputs are shifted and viewed as alpha and beta in a beta distribution. For illustrative purposes I generate some data and aim to maximize the average probability. The following code produces some gradients:

import tensorflow as tf
import tensorflow_probability as tfp

# a dummy model that will take in 3-dimensional data and output 2-dimensional data
model = tf.keras.Sequential()
model.add(tf.keras.layers.Dense(2,activation='relu'))

# data to feed into the model
input_data = tensorflow.random.normal((32,3))
# data to evaluate the distributions on
dist_data = tensorflow.random.uniform((32,1))/2.0+0.25

with tensorflow.GradientTape() as tape:
  # generate beta distributions from output of the network
  output = model(input_data)
  alpha = output[:,0,None]+1
  beta = output[:,1,None]+1
  dists = tfp.distributions.Beta(alpha,beta)

  # compute average probabilty
  dist_output = tensorflow.reduce_mean(dists.prob(dist_data))
grads = tape.gradient(dist_output,model.trainable_variables)
print(grads)

Output

[<tf.Tensor: shape=(3, 2), dtype=float32, numpy=
array([[-0.02362663,  0.05741577],
       [ 0.02077126,  0.03589955],
       [ 0.04922515, -0.0180204 ]], dtype=float32)>, <tf.Tensor: shape=(2,), dtype=float32, numpy=array([0.01512624, 0.08244684], dtype=float32)>]

Now if I instead target the CDF rather than the PDF I get None for gradients:

with tensorflow.GradientTape() as tape:
  # generate beta distributions from output of the network
  output = model(input_data)
  alpha = output[:,0,None]+1
  beta = output[:,1,None]+1
  dists = tfp.distributions.Beta(alpha,beta)

  # compute average cdf
  dist_output = tensorflow.reduce_mean(dists.cdf(dist_data))
grads = tape.gradient(dist_output,model.trainable_variables)
print(grads)

Output

[None, None]

I went to the underlying tensorflow_probability code and it appears that tensorflow_probability uses betainc from tensorflow.python.ops.gen_math_ops and that the gradients are not evaluated there. This can be seen by replacing the CDF above with a direct all to the incomplete beta function (grads still return None).

with tensorflow.GradientTape() as tape:
  # generate beta distributions from output of the network
  output = model(input_data)
  alpha = output[:,0,None]+1
  beta = output[:,1,None]+1
  # compute average cdf
  dist_output = tensorflow.reduce_mean(tensorflow.math.betainc(alpha,beta,dist_data))
grads = tape.gradient(dist_output,model.trainable_variables)
print(grads)

Output

[None, None]

Similarly, the PDF code above calls lbeta which also doesn't seem to be considered by the gradient, though in this case 0 is returned instead of None

with tensorflow.GradientTape() as tape:
  # generate beta distributions from output of the network
  output = model(input_data)
  alpha = output[:,0,None]+1
  beta = output[:,1,None]+1

  # compute average probabilty
  norm_output = tensorflow.reduce_mean(tensorflow.math.lbeta(alpha,beta))
  
norm_grads = tape.gradient(norm_output,model.trainable_variables)
print(norm_grads)

Output

[<tf.Tensor: shape=(3, 2), dtype=float32, numpy=
array([[0., 0.],
       [0., 0.],
       [0., 0.]], dtype=float32)>, <tf.Tensor: shape=(2,), dtype=float32, numpy=array([0., 0.], dtype=float32)>]

The net result of this is that if we have a complicated objective function involving both the PDF and the CDF, then we get non-zero gradients, but they are not correct as certain terms are not differentiated.

Question: What are my options for computing these gradients in the tensorflow framework?

0 Answers
Related