Unexpected action distribution for custom RL environment

Viewed 99

I am working on creating a custom environment and training a RL agent on it.

I am using stable-baselines because it seems to implement all the latest RL algorithms, and seems to be as close to "plug and play" as possible (I'd like to concentrate on creating the environment and reward function rather that the implementation details of the model itself)

My environment has an action space of size 127, and interprets it as a one-hot vector: Taking the index of the highest value in the vector as an input value. For debugging, I create a bar chart, showing how many times each value has been "called"

Before training, I would expect the graph to show a roughly uniform distribution of "events": uniform bar chart

but instead the "events" in the lower end of the action spec are massively more likely than the others: enter image description here

I have created a colab to explain and reproduce the issue

I asked this question in a github issue, but they recommended I post the question here

1 Answers

The model.predict(obs) clips each action to the range [-1, 1] (because that is how you defined your action space). Your array of action values thus looks something like

print(action)
# [-0.2476,  0.7068,  1.,          -1.,           1.,           1., 
#   0.1005,  -0.937,   -1. , ...]

That is, all actions that were larger than 1 are truncated / clipped to 1, and thus there are multiple maximum actions. In your environment you compute the numpy argmax pitch = np.argmax(action), which returns the index of the first maximum value rather than a randomly chosen one (if there are multiple maxima).

You could choose a "random argmax" as follows.

max_indeces = np.where(action == action.max())[0]
any_argmax = np.random.choice(max_indeces)

I changed your env accordingly here.

Related