Understanding multi agent learning in OpenAI gym and stable-baselines

Viewed 450

I was trying out developing multiagent reinforcement learning model using OpenAI stable baselines and gym as explained in this article.

I am confused about how do we specify opponent agents. It seems that opponents are passed to environment, as in case of agent2 below:

class ConnectFourGym:
    def __init__(self, agent2="random"):
        ks_env = make("connectx", debug=True)
        self.env = ks_env.train([None, agent2])

The ks_env.train() method seems to be the one from kaggle_environments.Environment:

def train(self, agents=[]):
    """
    Setup a lightweight training environment for a single agent.
    Note: This is designed to be a lightweight starting point which can
          be integrated with other frameworks (i.e. gym, stable-baselines).
          The reward returned by the "step" function here is a diff between the
          current and the previous step.
    Example:
        env = make("tictactoe")
        # Training agent in first position (player 1) against the default random agent.
        trainer = env.train([None, "random"])

Q1. However I got confused. Why does ConnectFourGym.__init__() calls train() method? That is why environment should do training? I feel, train() should be part of the model: above article uses PPO algorithm which contains train() method. This PPO.train() gets called when we call PPO.learn() which makes sense.

Q2. But then, reading PPO.learn()'s code, I dont see how it trains current agent against multiple opponent agents. Shouldnt model algo do this? Am reading it wrong? Or model is unaware of number of agents, its just known to environment and thats why environment contains train()? In that case, why does we have explicit Environment.train() method? Environment will return rewards as per multiple agent behavior and model will learn from that.

Or am completely screwed in basic concepts? Can somoene help me out?

0 Answers
Related