Policies and Value Functions
An MDP gives the rules of the problem. A policy chooses actions within those rules. A value function measures the expected return when following a policy.
Policies: what will the agent do?
Section titled “Policies: what will the agent do?”A deterministic policy chooses one action for each nonterminal state:
For our gridworld, consider the policy that chooses right in Start and Hall and up in B, C, and D. This completely specifies its behavior in every nonterminal cell.
A stochastic policy assigns probabilities to actions:
For example, in Hall it might choose right with probability 0.8 and down with probability 0.2. This randomness belongs to the agent’s choice. Slippery movement belongs to the environment’s response. Either or both can be random.
State value and action value
Section titled “State value and action value”The state-value function asks: “Starting in this state, what return do I expect if I follow policy ?”
The action-value function asks: “What if I take this particular action first, then follow ?”
The expectation averages over any randomness in actions, transitions, and rewards. A return describes one realized trajectory; a value is an expectation over the possible trajectories.
| Quantity | First action | Later actions |
|---|---|---|
| Chosen by . | Follow . | |
| Fix it to . | Follow . |
Values in the gridworld
Section titled “Values in the gridworld”Under the deterministic policy above and :
| State | Rewards still to come | State value |
|---|---|---|
| Goal | None: the episode has ended. | |
| Hall | ||
| Start | , then |
Entering Hall pays −1, but Hall has value 5 because of the reward that comes next. A different policy can give the same cell a different value: a policy that repeatedly hits a wall in Hall never collects the goal reward.
Bellman expectation: look ahead one step
Section titled “Bellman expectation: look ahead one step”The return satisfies:
Taking expectations under a policy gives the Bellman expectation equation:
In words: value here = expected next reward + discounted value of where you land.
For our deterministic route:
For a finite MDP, writing out both sources of randomness gives:
The inner sum averages the environment’s possible outcomes for an action. The outer sum averages over actions chosen by the policy. Neither sum means “choose the best”; we are evaluating the given policy.
Connect state and action values
Section titled “Connect state and action values”Fix the first action and average over the environment’s response:
Average those action values using the policy to recover state value:
Substituting this relationship at the next state gives the action-value Bellman expectation equation:
At a terminal next state, the future contribution is zero; the inner action average is omitted or represented using a zero-valued dummy action.
Preview: iterative policy evaluation
Section titled “Preview: iterative policy evaluation”For a finite discounted MDP with a known model, one way to evaluate a policy is to initialize estimates and repeatedly apply Bellman updates:
Here counts update sweeps, not environment time steps. Keep terminal values at zero. For , these synchronous updates converge to ; a practical implementation stops when changes fall below a chosen tolerance.
flowchart TD
accTitle: Iterative policy evaluation
accDescr: Initialize value estimates, apply a sweep of Bellman updates, and repeat until changes are small enough.
initialize["Initialize estimates; terminal values = 0"] --> sweep["Apply a Bellman update to every nonterminal state"]
sweep --> stable{"Are value changes small enough?"}
stable -->|"No"| sweep
stable -->|"Yes"| result["Estimated values for this policy"]
This is a preview of dynamic programming. The key distinction here is between the value we want and the estimates used to compute it.
Reference
Section titled “Reference”Sutton, R. S. & Barto, A. G. Reinforcement Learning: An Introduction, 2nd ed., chapter 3, section 3.5. The explanations and small worked examples here are written for these notes.
Previous: 3.4 Unified Notation for Episodic and Continuing Tasks · Next: 3.6 Optimal Policies and Optimal Value Functions