Recently I am doing research on multi-robot coverage. My goal is to optimize a parameterized policy, generate a transition matrix from each current state to their next states and make the stationary distribution on every state as uniform as possible.
This can be seen as maximizing the entropy over states. In this case, maximizing the entropy is equivalent to minimizing the KL divergence between a stationary distribution generated by the policy and a uniform distribution. However, my objective function is more complicated since it consists of other regularization events, thus the loss function is unknown.
There are two questions that remain unclear to me:
def forward(self):
# Get transition matrix
transitionMatrix = torch.stack([self.policy(self.policy.generateStateVector(i))
for i in range(self.policy.numS)]).to(self.policy.device).float()
# Calculate stationary distribution
PIE = torch.linalg.inv(transitionMatrix
- torch.eye(self.policy.numS).to(self.policy.device)
+ torch.ones(self.policy.numS, self.policy.numS).to(self.policy.device)).float()
stationaryDistribution = torch.matmul(torch.ones(1, self.policy.numS).to(self.policy.device), PIE)[0]
return stationaryDistribution, transitionMatrix
The code related to the first question is shown above. Since I designed a central optimizer to update the parameters of each robot in a receding horizon, the whole pipeline has to be separated into two parts of NN and matrix transformation, respectively. I am planning to feed the one-hot vector of each state into the NN, getting the transition probability vectors of states and stack them, thus obtaining the transition matrix. Can the gradient regarding to NN parameters be correctly back-propagated in this way?
obj = self.gamma_1 * H_overall + self.gamma_2 * CE_overall - self.gamma_3 * H_P
# Update parameters
self.policy.optimizer.zero_grad()
obj.backward()
self.policy.optimizer.step()
The code related to the second question is shown above. obj.backward() computes the gradient regarding to NN parameters of the current objective value. Since the optimizer.step() updates the parameters in a subtracting way, making obj=-obj should make θ=θ+grad * lr, thus maximizing the objective function. However, my intuition tells me that there must be some mistakes. Can anybody tell me whether it is correct or not?
Big thanks to all of you!