I am playing around with keras and reinforced learning. I created myself a sort of synthetic input data which looks like this:
The loss function seems to decrease its value within every episode, but then starts off again from random value with the beginning of every next minibatch. Looks like the agent forgets everything with every new episode.
The loss function looks like this (2 separate examples, each has 8 episodes, every batch has 100 samples)
What am I missing? What is this a symptom of?
