RL / LEARNING LAB Runs in your browser

04 / DYNAMIC PROGRAMMING

Small world. Big ideas.

Watch values propagate. See a policy improve.
Connect every update to the Python that makes it happen.

THE QUESTION

With a known model,
how do we find the best action?

Read the chapter ↗

Value landscape

Value update

Select a state to inspect its Bellman backup.

Inspected state★ TerminalArrows: greedy from stored Q
SWEEPS0
LAST SWEEP Δ—
POLICY UPDATES0

Ready. Start with one state update.

Getting closer

Δ by completed sweep · log₁₀(1 + Δ)

Run a sweep to see how much the values change.

TRY THIS

Evaluate, then optimize.

In the 2 × 2 world, run policy evaluation with γ = 1. State 0 approaches −8. Switch to value iteration: it reaches −2. Why does choosing an action beat averaging over actions?

Take the lab with you.

Host the built site, then paste this HTML into any page that allows iframes. All computation stays in the visitor’s browser.