I am trying to do a job shop scheduling which is solved with a reinfrocement q learning agent. This is what i got right now: https://git.uni-wuppertal.de/1523811/tmp
Unfortunately the agent does not learn well. I think I have to overthink the state representation. I have in that case 15 actions (1 for each Job) and also 15 Actions to place idle time (1 for each Job). In sum its 30 actions. But what is the state where the agent is? I dont know how to represent it.