Tensorflow 2 differentiate through optimization path?

Viewed 119

I am trying to compute "gradients through gradients" for a paper (MAML, by C.Finn et al.) in Tensorflow 2 with Keras backend. Thus, we start at some initial weights, compute K gradient update steps, and want to backpropagate through our initial weights. The code sample belows illustrates what I want to achieve, but unfortunately does not work.

optimizer = tf.keras.SGD()
initial_weights = model.trainable_variables

with tf.GradientTape() as mt:
    for gradient_steps in range(10):
        with tf.GradientTape() as t:
            loss = loss_function(y_train, model(x_train))
        grads = t.gradient(loss, model.trainable_variables)
        optimizer.apply_gradients(zip(grads, model.trainable_variables))
    test_loss = loss_function(y_test, model(x_test))
mt.gradient(test_loss, initial_weights)

Does anyone know how to differentiate through the initialization? Any help would be greatly appreciated!

0 Answers
Related