Need help for reward function in reinforcement learning

Viewed 75

I've created a RL to trade on an artificial custom financial asset (Complete Code). This is my dataframe (environment) made of 'Close' price and 'Volume':

loses = []
volumes = []
for i in range(0,16):
    for inc in range(0,30):
        closes.append(1 + 0.00005 * inc)
        volumes.append(2 + 0.00008 * inc)
    for dec in range(0,30):
        closes.append(1.00145 - 0.00005 * dec)
        volumes.append(2.00240 - 0.00008 * dec)

raw_df = pd.DataFrame(zip(closes, volumes), columns=['close','volume'])

I'm making my data stationary using differentiation (df - df.shift(1)) There are three actions: Sell, Buy and Hold.

And this is the returned observation after each step: 'close', 'volume', trade_length, total_episode_profit, current_profit, current_action (trading, watching)

There is an open and close trade cost equal to 1, and there is a 0.5 penalty to watch market and do nothing, holding reward is equal to close[-1] - close[-2] and sell reward is equal to total profit or loss of trading position.

And here is my NN structure:

model = Sequential()
        model.add(Dense(10, activation='tanh', input_shape=(env.df_ep.shape[1] + 4,)))
        model.add(Dropout(0.2))
        model.add(Dense(8))
        model.add(Dropout(0.2))
        model.add(Dense(env.ACTION_SPACE_SIZE, activation='linear'))
        model.compile(loss='mse', optimizer=adam_v2.Adam(learning_rate=0.001), metrics=['accuracy'])

The problem is after lots of episodes (about 6000) RL stops to learn and its just open a trade at the first and hold it till the end! But this is really a simple financial asset and a simple environment, it's not a real financial asset and I think it should learn it easily. I guess that the problem is with my reward function.

Here are some photos of episodes:

enter image description here

enter image description here

0 Answers
Related