Skip to content

Getting started

Chapters are ordered by difficulty, not by history. Each one assumes the chapters above it and nothing below it, so reading top to bottom always works.

ChapterWhat it adds
FoundationsThe three historical threads that became modern RL.
Multi-armed banditsExploration vs. exploitation, with no state.
Markov decision processesState, transitions, value functions, Bellman equations.
Dynamic programmingExact solutions when the model is known.
Monte Carlo methodsFirst model-free learning, from complete episodes.
Temporal-difference learningBootstrapping: TD(0), SARSA, Q-learning.
n-step and eligibility tracesThe axis connecting Monte Carlo and TD.
Planning and learningLearned models and simulated experience.
Function approximationDropping the lookup table.
Deep Q-networksNeural value approximation at scale.
Policy gradient methodsOptimising the policy directly.
Advanced policy optimizationActor-critic, trust regions, PPO.

Every topic page has six required sections — learning objective, intuition, algorithm and equations, an interactive demo, what to observe, and personal takeaways. Worked examples, common mistakes, experiments, and references appear when they earn their place.

Each demo runs as its own full-screen page and is also embedded in the topic page it belongs to. They are deterministic: every demo takes a seed, and the same seed always reproduces the same run. If something surprising happens, note the seed and it can be replayed exactly.

The simulation code is separated from the drawing code, so the same core that animates a demo also runs headlessly for the comparison charts.

Terminal window
npm install
npm run dev # local preview
npm run check # TypeScript and component diagnostics
npm run build # production build, including internal-link validation

Node is pinned in .nvmrc. Pushing to main builds and deploys to GitHub Pages automatically.