Concepts / Modern Reinforcement Learning

Modern Reinforcement Learning

Samuel's checkers player combined a search role with a learning role.

  • Programming

A Checkers Program with Two Jobs

Arthur Samuel's checkers-playing programs are historically important because they combined two different kinds of work. The program had to search through possible play efficiently, and it had to learn from play over time. Samuel developed these programs in work published in 1959 and 1967. Their importance is not that they represent every detail of modern reinforcement learning, but that they provide an early connection to ideas that later became central to that field.

The two methods associated with Samuel's checkers players were heuristic search for efficient exploration of possibilities and temporal-difference learning for learning from experience.

Separating Search from Learning

exploreshelps chooseprovides feedback over timeupdatesinforms later searchesCheckers positionHeuristic searchefficient explorationSelected movePlay experienceTemporal-differencelearninglearning over timeLearned knowledgeused in later searches
How does the program use heuristic search to choose a move, while its learning mechanism changes the knowledge used in future searches?

Search and learning answer different questions. Search asks how the program should explore possible play now. Learning asks how the program should improve its knowledge through play over time. In Samuel's work, heuristic search addressed search efficiency. Temporal-difference learning provided the learning approach associated with the programs.

A Conceptual Trace through Samuel's Program

Tracing one learning cycle

Explain what happens conceptually when a checkers-playing program encounters a position, selects a move, and later learns from its experience.

1. Encounter a position: The program begins with a checkers position and needs to decide how to proceed.

2. Search efficiently: Its search role explores possible play. Heuristic search is important because it improves the efficiency of that exploration.

3. Select a move: The search process supports a move choice. This is the program's immediate decision-making role.

4. Learn from play: Over time, the program uses temporal-difference learning as its learning approach. This changes the knowledge available for later play.

5. Connect the roles: The updated knowledge can matter to future searches, so the program's later decisions are connected to what it learned from earlier play.

The search role chooses how to explore possibilities efficiently, while the learning role changes what the program can use in future searches.

This trace is a conceptual teaching example, not a reconstruction of an undocumented implementation detail. Its purpose is to preserve the historical distinction in Samuel's work: searching is the mechanism for exploring possibilities, and learning is the mechanism for improving over time.

From Samuel to Reinforcement Learning

Samuel's work is considered a precursor to modern reinforcement learning because the checkers players used what is now called temporal-difference learning. The historical wording matters: the programs belong to an earlier period, while the modern name identifies the learning approach in terms used today. Samuel's programs therefore provide a connection between early game-playing research and the later field of reinforcement learning.

  • The program's task was to learn to play checkers.
  • Heuristic search improved the efficiency of exploring possible play.
  • Temporal-difference learning supplied the learning approach associated with the work.
  • The combination of search and learning makes the programs useful historical case studies.
  • The programs are precursors to modern reinforcement learning, not a complete description of all modern reinforcement-learning practice.

Three Phases of Multi-Agent Research

adds context and neural networksadds brain-related constraintsPhase 1non-associative learningautomataPhase 2associative learning andneural networksPhase 3neuroscience-informedlearning
What changed from one research phase to the next, and how did the focus of multi-agent reinforcement learning develop over time?
PhaseMain focusWhat it added
FirstNon-associative learning automataLearning in bandit, team, and game problems
SecondAssociative or contextual learningContext, stochastic learning automata, and artificial neural networks
ThirdNeuroscience-informed learningAttention to synaptic plasticity and constraints suggested by the brain

A broad historical sequence in research on teams of reinforcement learning agents

The three phases should be read as a widening sequence of questions. Early research asked how learning automata could behave in bandit, team, and game problems. Later research asked how a learner could use context and how learning units could form artificial neural networks. A subsequent phase asked how learning rules could better reflect findings about the brain.

Context Changes What Learning Means

Non-associative learning automataAssociative or contextual reinforcement learning
The first historical phase focused on learning behavior in bandit, team, and game problems.The second phase extended learning automata to the associative, or contextual, case.
The studies did not address the contextual bandit case.The learner is described as using an explicitly associative context.
The focus is on learning behavior without the contextual association emphasized in the later phase.The work connected stochastic learning automata with artificial neural networks.

The word non-associative marks the boundary of the first phase. Those studies investigated learning behavior, including behavior in teams and games, but did not address the contextual bandit case. The second phase added the associative or contextual case: researchers studied how a learner could use context while learning.

In the second phase, Barto, Sutton, and Brouwer, along with Barto and Sutton, experimented with associative stochastic learning automata in single-layer artificial neural networks. These networks received one global reinforcement signal, and their neuron-like learning elements were called associative search elements, or ASEs.

Why A R - P Mattered

Barto and Anandan introduced the associative reward-penalty algorithm, usually abbreviated A R - P, in 1985. It was a more sophisticated associative reinforcement learning algorithm. Its importance came from connecting several lines of work: stochastic learning automata, pattern classification, associative reinforcement learning, and artificial neural networks.

The source describes A R - P historically rather than as a complete implementation recipe. Its importance is the connection it made between context-sensitive reinforcement learning and neural-network learning. Teams of A R - P units were connected into multi-layer neural networks, and reported results showed that these teams could learn nonlinear functions, including XOR, while receiving a globally broadcast reinforcement signal.

Later, Williams mathematically analyzed and broadened this class of learning rules and showed in 1992 that a special case of A R - P is a REINFORCE algorithm.

Neuroscience Shapes the Third Phase

The third phase was influenced by growing neuroscience support for the possibility that this kind of learning occurs in the brain. Researchers began paying more attention to synaptic plasticity and to other constraints suggested by neuroscience, rather than treating a learning rule only as an abstract computational procedure.

  • STDP, or spike-timing-dependent plasticity
  • Dopamine
  • Reward-modulated STDP
  • Synaptic plasticity
  • Other biological constraints suggested by neuroscience

Common Historical Mistakes

  • Treating heuristic search and temporal-difference learning as the same technique.

    The source assigns search efficiency to heuristic search and assigns learning over time to temporal-difference learning.

    Fix: Describe heuristic search as the search role and temporal-difference learning as the learning role.

  • Describing Samuel's programs as a complete modern reinforcement-learning system.

    The source calls them a precursor and an early connection, not a complete account of modern reinforcement learning.

    Fix: Explain that Samuel's programs used an approach now called temporal-difference learning and are historically significant precursors.

  • Calling the first phase contextual reinforcement learning.

    The source specifically states that the first-phase studies did not address the contextual bandit case.

    Fix: Reserve associative or contextual learning for the second phase.

  • Reducing the third phase to neural networks alone.

    Its defining influence was growing neuroscience support and attention to synaptic plasticity and biological constraints.

    Fix: Connect the third phase to STDP, dopamine, reward-modulated STDP, synaptic plasticity, and related neuroscience constraints.

Check Your Understanding

MEDIUM

A learner says: Samuel's program used heuristic search to learn from experience, and the second historical phase was non-associative because it studied learning automata. Correct both parts of the statement.

Hints
  • Ask which method addressed search efficiency and which method provided learning over time.
  • Recall what the second phase added to learning automata.

Answer

Correct the two-part historical statement.

First correction: Heuristic search addressed efficient exploration of possible play. Temporal-difference learning provided the learning approach associated with Samuel's programs.

Second correction: The second phase extended learning automata to associative or contextual learning and connected them with artificial neural networks.

A corrected statement is: Samuel's program used heuristic search to make its search more efficient and temporal-difference learning to learn over time; the second phase added associative or contextual learning and connected learning automata with artificial neural networks.

Historical Takeaways

  1. Samuel's checkers players combined heuristic search with temporal-difference learning.
  2. Search explored possibilities efficiently; learning changed the knowledge available for later play.
  3. Samuel's programs are precursors to modern reinforcement learning because they used an approach now called temporal-difference learning.
  4. Research on teams of reinforcement learning agents moved from non-associative learning automata, to associative or contextual learning with neural networks, and then toward neuroscience-informed learning.
  5. A R - P strengthened the connection among stochastic learning automata, associative reinforcement learning, pattern classification, and artificial neural networks, while neuroscience directed later research toward synaptic plasticity and related biological constraints.

Key Takeaways

  • Samuel's checkers work is best understood by separating its search role from its learning role.
  • Heuristic search improved the efficiency of exploring possible play, while temporal-difference learning supported learning over time.
  • The history of teams of reinforcement learning agents is commonly described in three phases: non-associative learning automata, associative or contextual learning with neural networks, and neuroscience-informed research.
  • The associative reward-penalty algorithm connected several research traditions and was later shown to include a special case of REINFORCE.
  • Neuroscience influenced the third phase through attention to STDP, dopamine, reward-modulated STDP, synaptic plasticity, and other biological constraints.