Skip to content

Temporal-difference learning

Bootstrap: update a value estimate from another estimate, without waiting for the episode to end. The core idea of modern RL.

  • The TD(0) update and the TD error δt\delta_t
  • TD vs. Monte Carlo: bias, variance, and speed
  • SARSA — on-policy control
  • Q-learning — off-policy control
  • Expected SARSA and maximisation bias / double Q-learning

Everything above this chapter in the sidebar.