Temporal-difference learning
Learning objective
Section titled “Learning objective”Bootstrap: update a value estimate from another estimate, without waiting for the episode to end. The core idea of modern RL.
Planned scope
Section titled “Planned scope”- The TD(0) update and the TD error
- TD vs. Monte Carlo: bias, variance, and speed
- SARSA — on-policy control
- Q-learning — off-policy control
- Expected SARSA and maximisation bias / double Q-learning
Prerequisites
Section titled “Prerequisites”Everything above this chapter in the sidebar.