I'm trying to do this
np.random.choice(a, 1, p=p)
where len(a) != len(p)
Could you point me in the direction where to look for how to resize the probability distribution "p"? The idea is to keep the same distribution but over a different number of variables.
EDIT: Basically this (https://en.wikipedia.org/wiki/Scale_parameter) but with discrete variable.
I think that the interpolation is the way to go as suggested by Ryan Sander. I am using a neural network to output the policy distribution over an environment action space. I'm trying to train the network on multiple environments with different action space sizes. For example the network is outputting the distribution over a action space of size 6 (actions [0,1,2,...5], 6 numbers summing up to 1) and I'm trying to sample this distribution over an action space of size 9. Or the other way around.
The problem with interpolation is that the values that I get do not sum up to 1. If I do softmax on those values, the distribution that I get does not have the same(ish) shape as the original.