Learning MDP but this Lotto question has me stumped

Viewed 18

I'm trying to understand MDP more as I get into reinforcement learning but this question has me stumped. if anyone has any clues as to how I would be best to go about this:

Assume that Bob has $20 to buy lotto tickets. Bob will continue to buy lotto tickets every Saturday until he loses all the money or as soon as he has $40 or more money. On each Saturday, if Bob still has money, Bob will choose to buy either of two different lotto tickets, i.e., ticket A or ticket B (Bob can only buy at most one ticket on any Saturday). Ticket A costs $10. With ticket A, Bob has the probability of 0.2% to earn $20 and 0.2% to learn $30 and $0 otherwise. Ticket B costs $20. With ticket B, Bob has the probability of 0.2% to earn $30 and 0.2% to earn $40 and $0 otherwise. Provide a Markov decision process (MDP) model of the above problem. In your MDP model, clearly describe the state space, action space, rewards and transition probabilities. Assume that the discount factor γ = 1.0. In this problem, Bob has the goal to double his wealth (i.e., from $20 to $40 or more). Rewards should be defined properly to match this goal.

0 Answers
Related