Policy evaluation
Keep π fixed. Replace each V(s) with Σ π(a|s)Q(s,a), then repeat until a complete sweep has Δ < θ. Finally recompute all Q from the converged V.
value_prediction(env, pi, …)04 / DYNAMIC PROGRAMMING
Watch values propagate. See a policy improve.
Connect every update to the Python that makes it happen.
Select a state to inspect its Bellman backup.
Ready. Start with one state update.
Run a sweep to see how much the values change.
TRY THIS
In the 2 × 2 world, run policy evaluation with γ = 1. State 0 approaches −8. Switch to value iteration: it reaches −2. Why does choosing an action beat averaging over actions?
THE LOGICAL STRUCTURE
Dynamic programming uses the transition model to look ahead at every possible outcome. No episodes are sampled in this lab.
env.unwrapped.PFor every (state, action), read probability, next state, reward, and done.
Q(s,a) = Σ p(r + γV(s′))Use zero future value when done is true. Otherwise read the current V.
V(s) ← average or maxAverage with π to evaluate. Take the maximum to optimize.
Keep π fixed. Replace each V(s) with Σ π(a|s)Q(s,a), then repeat until a complete sweep has Δ < θ. Finally recompute all Q from the converged V.
value_prediction(env, pi, …)Replace each V(s) with max Q(s,a). Repeat until Δ < θ, then return a deterministic greedy policy from the last sweep’s Q table.
value_iteration(env, …)Start with a uniform random policy. Evaluate it to tolerance, improve all states with argmax Q, and repeat. Stop when the policy no longer changes.
evaluate → improve → repeatYour code assigns V[state] = new_v immediately. Later states in the same sweep can read that new value. “Step state” exposes this ordering; it is not a synchronous update using a frozen copy of V.
Δ is the largest absolute change to V in one sweep. It is a stopping criterion, not a measurement of error against the true optimal value. The threshold applies only after every state has been updated.
np.argmax takes the first maximum. Action order is left (0), up (1), right (2), down (3). The browser follows that same order, including temporary ties early in learning.
V estimates return; π chooses actions. Evaluation changes only V. Policy improvement changes π. Value iteration combines a greedy choice with each value backup; its displayed arrows preview the eventual greedy policy.