Bain: learning by “groping and experiment”
Woodworth traces the idea of trial-and-error learning back to Alexander Bain’s discussion in the 1850s.
Reinforcement learning did not begin as one neat branch of machine learning. For more than a century, three groups worked on different versions of the same question without yet sharing a language:
Modern RL appeared when those three threads were finally braided together.
This is the oldest thread. An action followed by a satisfying result becomes more likely; an action followed by discomfort becomes less likely. It sounds simple, but it creates the defining RL dilemma: the learner must act to discover which actions are good, and every experiment has a real consequence.
This thread gives RL its agent, actions, rewards, and the tension between exploration and exploitation.
Control theory approached the problem from mathematics rather than psychology. Given a model of a changing system, which sequence of decisions produces the best total outcome? Richard Bellman’s answer was to break the problem into smaller problems and assign each situation a value.
This thread gives RL Markov decision processes, value functions, Bellman equations, and dynamic programming.
The third thread learns by comparing two successive predictions. You do not have to wait until a chess game—or a life—ends before updating an earlier estimate. If the next moment looks better than expected, revise upward now; if it looks worse, revise downward.
This thread gives RL bootstrapping: improving an estimate using another estimate before the final outcome is known.
Use the filters to isolate one lineage, then return to All threads and watch the ideas begin to cross. The brighter cards mark the moments that changed the direction of the field.
Showing all 30 moments
Woodworth traces the idea of trial-and-error learning back to Alexander Bain’s discussion in the 1850s.
The British ethologist uses the term explicitly to describe what he observes in animal behaviour.
Responses followed by satisfaction become more likely to recur; those followed by discomfort, less likely. This is the first succinct statement of the principle.
The word first appears in this context in the English translation of Pavlov’s monograph on conditioned reflexes.
Possibly the earliest electro-mechanical learner: it finds its way through a simple maze and remembers the path in switch settings.
A report sketches a machine that makes random choices where action is undetermined, cancelling them on pain and fixing them on pleasure.
Shannon suggests a chess program could improve online by modifying its evaluation function—the seed of Samuel’s method.
Mechanical creatures learn simple behaviours; Shannon’s maze-running mouse finds a route by trial and error while the maze remembers through magnets and relays.
Minsky describes Stochastic Neural-Analog Reinforcement Calculators, while Farley and Clark simulate a network that learns by trial and error.
A stimulus paired with a primary reinforcer can itself become predictive—a psychological idea with consequences for artificial learning systems.
Bellman defines the value function and Bellman equation, then introduces the discrete stochastic case: Markov decision processes.
The first learning method to implement temporal-difference ideas is developed with no reference to Minsky or animal learning.
A method for solving Markov decision processes that remains an essential element of modern reinforcement-learning theory.
How should credit for success be distributed among the many decisions that produced it?
Michie’s matchbox engine learns tic-tac-toe with coloured beads. Andreae’s STeLLA learns by trial and error with an internal model of the world.
Waltz and Fu—followed by Mendel, Fu, and others—use the terms for engineering applications of trial-and-error learning.
Michie and Chambers build a controller that learns from a failure signal alone—an early RL task under incomplete knowledge.
Klopf argues that the drive to get a result from the environment was being lost as research turned toward supervised learning.
Widrow, Gupta, and Maitra modify LMS to learn from success and failure instead of examples. Learning automata and the k-armed bandit join the thread.
Selection-based classifier systems lead to a bucket-brigade credit-assignment algorithm with a family resemblance to temporal difference.
What we now call tabular TD(0) appears inside an adaptive controller for Markov decision processes.
Sutton develops learning rules driven by changes between temporally successive predictions, with links back to animal learning.
Sutton and Barto build a psychological model of classical conditioning on temporal-difference learning.
Barto, Sutton, and Anderson combine temporal-difference learning with trial and error, then return to the pole-balancing problem.
An explicit case for joining dynamic programming with learning methods—and for connecting both to neural and cognitive mechanisms.
Sutton treats temporal-difference learning as a general prediction method, introduces TD(λ), and proves convergence properties.
Watkins integrates trial-and-error learning, optimal control, and temporal difference inside the Markov decision process formalism.
Tesauro combines temporal-difference learning with a neural network for backgammon’s enormous state space, eventually reaching elite human play.
Researchers find a striking similarity between temporal-difference errors and dopamine-neuron activity, opening a new thread in neuroscience.
Bertsekas and Tsitsiklis develop the relationship between dynamic programming and artificial neural networks that Watkins helped open.
Watkins’s Q-learning is the hinge of the story. It learns by trial and error, uses the value-based structure of optimal control, and updates estimates by a temporal difference. After 1989, “reinforcement learning” increasingly names this combined field rather than any one of its ancestors.
The familiar agent–environment loop is what the braid looks like in compact form:
An agent observes a situation , chooses an action , and the environment returns a reward and a new situation. A policy chooses actions; a value estimate predicts long-term reward. Every chapter that follows changes one part of this loop or one way of learning from it.
| Historical thread | Question it contributes | Where it reappears |
|---|---|---|
| Trial and error | Which action should I try? | Multi-armed bandits and exploration |
| Optimal control | What are future outcomes worth now? | MDPs and dynamic programming |
| Temporal difference | How can one prediction improve another? | TD learning and eligibility traces |
| The braid | How can an agent learn control from experience? | Q-learning, DQNs, and actor–critic |
The next chapter removes state, delayed consequences, and almost everything else. What remains is the trial-and-error problem in its purest form: multi-armed bandits.