I'm a reading a book, "Reinforcement Learning" by S. Sutton and Andrew G. Barto
Sutton's Book explains: "For Monte Carlo policy iteration it is natural to alternate between evaluation and improvement on an episode-by-episode basis. After each episode, the observed returns are used for policy evaluation, and then the policy is improved at all the states visited in the episode. A complete simple algorithm along these lines, which we call Monte Carlo ES, for Monte Carlo with Exploring Starts, is given in pseudo code in the box on the next page"

However, in this pseudo code, it seems last two lines, which correspond to policy evaluation and policy improvement are done for each step of episode. According to Sutton's explanation, should't last two lines be moved 2 indentations to left so that policy evaluation and policy improvement are performed for each episode?