Skip to content

Markov decision processes

A Markov decision process (MDP) describes sequential decisions: the agent observes a state, chooses an action, and receives a reward and a next state. The goal is to choose a policy with high expected return.

MDP overview covering the agent–environment loop, states, actions, rewards, trajectories, policies, value functions, Bellman equations, and iterative policy evaluation.

The image is a quick map of the chapter. “Reward + future value” in its Bellman panel means expected reward plus discounted future value. The policy-evaluation panel previews methods developed in the next chapter.

The seven sections follow this order. They share a small gridworld introduced in section 3.1.

SectionWhat you will learn
3.1 The Agent–Environment InterfaceStates, actions, transition dynamics, and the Markov property.
3.2 Goals and RewardsExpress the task through rewards while distinguishing immediate feedback from long-term success.
3.3 Returns and EpisodesCalculate finite and discounted returns, including backward calculations and infinite sequences.
3.4 Unified Notation for Episodic and Continuing TasksUse an absorbing terminal state and zero future rewards to write one return formula for both task types.
3.5 Policies and Value FunctionsDefine policies, state and action values, and the Bellman expectation equations.
3.6 Optimal Policies and Optimal Value FunctionsCompare policies, derive Bellman optimality, and choose actions from optimal values.
3.7 Optimality and ApproximationUnderstand why optimality is a useful target and why practical agents use estimates and approximation.

Sutton, R. S. & Barto, A. G. Reinforcement Learning: An Introduction, 2nd ed., chapter 3: Finite Markov Decision Processes.