Skip to content

Three threads of reinforcement learning

Reinforcement learning did not begin as one neat branch of machine learning. For more than a century, three groups worked on different versions of the same question without yet sharing a language:

  • Psychologists asked how consequences change behaviour.
  • Control theorists asked how a sequence of decisions can maximise a long-term objective.
  • Prediction researchers asked how one estimate can learn from the next.

Modern RL appeared when those three threads were finally braided together.

This is the oldest thread. An action followed by a satisfying result becomes more likely; an action followed by discomfort becomes less likely. It sounds simple, but it creates the defining RL dilemma: the learner must act to discover which actions are good, and every experiment has a real consequence.

This thread gives RL its agent, actions, rewards, and the tension between exploration and exploitation.

Optimal control: what is this situation worth?

Section titled “Optimal control: what is this situation worth?”

Control theory approached the problem from mathematics rather than psychology. Given a model of a changing system, which sequence of decisions produces the best total outcome? Richard Bellman’s answer was to break the problem into smaller problems and assign each situation a value.

This thread gives RL Markov decision processes, value functions, Bellman equations, and dynamic programming.

Temporal difference: was my prediction too high or too low?

Section titled “Temporal difference: was my prediction too high or too low?”

The third thread learns by comparing two successive predictions. You do not have to wait until a chess game—or a life—ends before updating an earlier estimate. If the next moment looks better than expected, revise upward now; if it looks worse, revise downward.

This thread gives RL bootstrapping: improving an estimate using another estimate before the final outcome is known.

Use the filters to isolate one lineage, then return to All threads and watch the ideas begin to cross. The brighter cards mark the moments that changed the direction of the field.

Showing all 30 moments

  1. Trial and error

    Bain: learning by “groping and experiment”

    Woodworth traces the idea of trial-and-error learning back to Alexander Bain’s discussion in the 1850s.

  2. Trial and error

    Lloyd Morgan names “trial and error”

    The British ethologist uses the term explicitly to describe what he observes in animal behaviour.

  3. Trial and error

    “Reinforcement” enters the vocabulary

    The word first appears in this context in the English translation of Pavlov’s monograph on conditioned reflexes.

  4. Trial and error

    Ross builds a maze-learning machine

    Possibly the earliest electro-mechanical learner: it finds its way through a simple maze and remembers the path in switch settings.

  5. Trial and error

    Turing’s “pleasure–pain system”

    A report sketches a machine that makes random choices where action is undetermined, cancelling them on pain and fixing them on pleasure.

  6. Temporal difference

    Shannon: a program that improves its own evaluation

    Shannon suggests a chess program could improve online by modifying its evaluation function—the seed of Samuel’s method.

  7. Trial and error

    Walter’s tortoise and Shannon’s Theseus

    Mechanical creatures learn simple behaviours; Shannon’s maze-running mouse finds a route by trial and error while the maze remembers through magnets and relays.

  8. Trial and error

    SNARCs and simulated neural learning

    Minsky describes Stochastic Neural-Analog Reinforcement Calculators, while Farley and Clark simulate a network that learns by trial and error.

  9. Temporal difference

    Minsky links secondary reinforcement to machines

    A stimulus paired with a primary reinforcer can itself become predictive—a psychological idea with consequences for artificial learning systems.

  10. Optimal control

    Howard’s policy iteration

    A method for solving Markov decision processes that remains an essential element of modern reinforcement-learning theory.

  11. Trial and error

    Minsky names the credit-assignment problem

    How should credit for success be distributed among the many decisions that produced it?

  12. Trial and error

    MENACE and STeLLA

    Michie’s matchbox engine learns tic-tac-toe with coloured beads. Andreae’s STeLLA learns by trial and error with an internal model of the world.

  13. Trial and error

    “Reinforcement learning” enters engineering

    Waltz and Fu—followed by Mendel, Fu, and others—use the terms for engineering applications of trial-and-error learning.

  14. Trial and error

    BOXES balances a pole

    Michie and Chambers build a controller that learns from a failure signal alone—an early RL task under incomplete knowledge.

  15. Trial and error

    Klopf revives the hedonic thread

    Klopf argues that the drive to get a result from the environment was being lost as research turned toward supervised learning.

  16. Trial and error

    Learning with a critic

    Widrow, Gupta, and Maitra modify LMS to learn from success and failure instead of examples. Learning automata and the k-armed bandit join the thread.

  17. Trial and error

    Holland’s adaptive systems

    Selection-based classifier systems lead to a bucket-brigade credit-assignment algorithm with a family resemblance to temporal difference.

  18. Temporal difference

    Learning from successive predictions

    Sutton develops learning rules driven by changes between temporally successive predictions, with links back to animal learning.

  19. Temporal difference

    A TD model of classical conditioning

    Sutton and Barto build a psychological model of classical conditioning on temporal-difference learning.

  20. Optimal control

    Werbos argues for interrelating DP and learning

    An explicit case for joining dynamic programming with learning methods—and for connecting both to neural and cognitive mechanisms.

  21. Optimal control

    Neurodynamic programming

    Bertsekas and Tsitsiklis develop the relationship between dynamic programming and artificial neural networks that Watkins helped open.

Watkins’s Q-learning is the hinge of the story. It learns by trial and error, uses the value-based structure of optimal control, and updates estimates by a temporal difference. After 1989, “reinforcement learning” increasingly names this combined field rather than any one of its ancestors.

The familiar agent–environment loop is what the braid looks like in compact form:

St→AtRt+1,St+1S_t \xrightarrow{A_t} R_{t+1}, S_{t+1}

An agent observes a situation StS_t, chooses an action AtA_t, and the environment returns a reward Rt+1R_{t+1} and a new situation. A policy chooses actions; a value estimate predicts long-term reward. Every chapter that follows changes one part of this loop or one way of learning from it.

The map of the book is hidden in the history

Section titled “The map of the book is hidden in the history”
Historical threadQuestion it contributesWhere it reappears
Trial and errorWhich action should I try?Multi-armed bandits and exploration
Optimal controlWhat are future outcomes worth now?MDPs and dynamic programming
Temporal differenceHow can one prediction improve another?TD learning and eligibility traces
The braidHow can an agent learn control from experience?Q-learning, DQNs, and actor–critic

The next chapter removes state, delayed consequences, and almost everything else. What remains is the trial-and-error problem in its purest form: multi-armed bandits.