I've created a RL to trade on an artificial custom financial asset (Complete Code). This is my dataframe (environment) made of 'Close' price and 'Volume':
loses = []
volumes = []
for i in range(0,16):
for inc in range(0,30):
closes.append(1 + 0.00005 * inc)
volumes.append(2 + 0.00008 * inc)
for dec in range(0,30):
closes.append(1.00145 - 0.00005 * dec)
volumes.append(2.00240 - 0.00008 * dec)
raw_df = pd.DataFrame(zip(closes, volumes), columns=['close','volume'])
I'm making my data stationary using differentiation (df - df.shift(1)) There are three actions: Sell, Buy and Hold.
And this is the returned observation after each step: 'close', 'volume', trade_length, total_episode_profit, current_profit, current_action (trading, watching)
There is an open and close trade cost equal to 1, and there is a 0.5 penalty to watch market and do nothing, holding reward is equal to close[-1] - close[-2] and sell reward is equal to total profit or loss of trading position.
And here is my NN structure:
model = Sequential()
model.add(Dense(10, activation='tanh', input_shape=(env.df_ep.shape[1] + 4,)))
model.add(Dropout(0.2))
model.add(Dense(8))
model.add(Dropout(0.2))
model.add(Dense(env.ACTION_SPACE_SIZE, activation='linear'))
model.compile(loss='mse', optimizer=adam_v2.Adam(learning_rate=0.001), metrics=['accuracy'])
The problem is after lots of episodes (about 6000) RL stops to learn and its just open a trade at the first and hold it till the end! But this is really a simple financial asset and a simple environment, it's not a real financial asset and I think it should learn it easily. I guess that the problem is with my reward function.
Here are some photos of episodes:

