tensorflow, compute gradients with respect to weights that come from two models (encoder, decoder)

Viewed 390

I have a encoder model and a decoder model (RNN). I want to compute the gradients and update the weights. I'm somewhat confused by what I've seen so far on the web. Which block is the best practice? Is there any difference between the two options? Gradients seems to converge faster in Block 1, I do not know why?

# BLOCK 1, in two operations
encoder_gradients,decoder_gradients = tape.gradient(loss,[encoder_model.trainable_variables,decoder_model.trainable_variables])
myoptimizer.apply_gradients(zip(encoder_gradients,encoder_model.trainable_variables))
myoptimizer.apply_gradients(zip(decoder_gradients,decoder_model.trainable_variables))
# BLOCK 2, in one operation
gradients = tape.gradient(loss,encoder_model.trainable_variables + decoder_model.trainable_variables)
myoptimizer.apply_gradients(zip(gradients,encoder_model.trainable_variables +
decoder_model.trainable_variables))
1 Answers

You can manually verify this.

First, let's simplify the model. Let the encoder and decoder both be a single dense layer. This is mostly for simplicity and you can print out the weights being applying the gradients, gradients and weights after applying the gradients.

import tensorflow as tf
import numpy as np
from copy import deepcopy

# create a simple model with one encoder and one decoder layer. 
class custom_net(tf.keras.Model):
    def __init__(self):
        super().__init__()
        self.encoder = tf.keras.layers.Dense(3, activation='relu')
        self.decoder = tf.keras.layers.Dense(3, activation='relu')
        
    def call(self, inp):
        return self.decoder(self.encoder(inp))

net = model()

# create dummy input/output
inp = np.random.randn(1,1)
gt = np.random.randn(3,1)

# set persistent to true since we will be accessing the gradient 2 times
with tf.GradientTape(persistent=True) as tape:
    out = custom_model(inp)
    loss = tf.keras.losses.mean_squared_error(gt, out)
    
# get the gradients as mentioned in the question
enc_grad, dec_grad = tape.gradient(loss,
                             [net.encoder.trainable_variables, 
                              net.decoder.trainable_variables])
gradients = tape.gradient(loss,
                          net.encoder.trainable_variables + net.decoder.trainable_variables)

First, let's use a stateless optimizer like SGD which updates the weights based on the following formula and compare it to the 2 approaches mentioned in the question.

new_weights = weights - learning_rate * gradients.

# Block 1

myoptimizer = tf.keras.optimizers.SGD(learning_rate=1)

# store weights before updating the weights based on the gradients 
old_enc_weights = deepcopy(net.encoder.get_weights())
old_dec_weights = deepcopy(net.decoder.get_weights())

myoptimizer.apply_gradients(zip(enc_grad, net.encoder.trainable_variables))
myoptimizer.apply_gradients(zip(dec_grad, net.decoder.trainable_variables))

# manually calculate the weights after gradient update
# since the learning rate is 1, new_weights = weights - grad
cal_enc_weights = []
for weights, grad in zip(old_enc_weights, enc_grad):
    cal_enc_weights.append(weights-grad)

cal_dec_weights = []
for weights, grad in zip(old_dec_weights, dec_grad):
    cal_dec_weights.append(weights-grad)    
    
for weights, man_calc_weight in zip(net.encoder.get_weights(), cal_enc_weights):
    print(np.linalg.norm(weights-man_calc_weight))

for weights, man_calc_weight in zip(net.decoder.get_weights(), cal_dec_weights):
    print(np.linalg.norm(weights-man_calc_weight))

# block 2 
old_weights = deepcopy(net.encoder.trainable_variables + net.decoder.trainable_variables)
myoptimizer.apply_gradients(zip(gradients, net.encoder.trainable_variables + \
                                net.decoder.trainable_variables))
cal_weights = []
for weight, grad in zip(old_weights, gradients):
    cal_weights.append(weight-grad) 
    
for weight, man_calc_weight in zip(net.encoder.trainable_variables + net.decoder.trainable_variables, cal_weights):
    print(np.linalg.norm(weight-man_calc_weight))   

You will see that both the methods update the weights in the exact same way.

I think you used an optimizer like Adam/RMSProp which is stateful. For such optimizers invoking apply_gradients will update the optimizer parameters based on the gradient value and sign. In the first case, the optimizer parameters are updated twice and in the second case only once.

I would stick to the second option if I were you, since you are performing just one step of optimization here.

Related