Concepts / Self-Play Learning Methods

Self-Play Learning Methods

TD-Gammon's importance came from learning strong play with little backgammon knowledge.

  • Programming

A Different Route to Strong Play

TD-Gammon became important for more than simply playing backgammon well. Earlier successful programs relied heavily on expert backgammon knowledge, carefully designed features, or training examples supplied by experts. TD-Gammon instead learned through self-play with essentially zero backgammon knowledge built in. That made its route to strong play the central achievement.

useslearns throughEarlier programsexpert knowledge andfeaturesExpert examplestraining movesTD-Gammonessentially zero backgammonknowledgeSelf-playgenerated experience
What was different about TD-Gammon's source of playing strength?

The key comparison is not simply handcrafted play versus neural networks. TD-Gammon combined a multilayer neural network with a learning process that generated its own experience through self-play, rather than depending on a large collection of expert moves.

One Self-Play Improvement Cycle

Self-play changed where the learning experience came from. TD-Gammon played backgammon against itself. The games produced positions and outcomes that the learning method could use to adjust the program's evaluation. Repeating this process allowed the program to learn from its own games instead of waiting for experts to provide every important move.

selectsextendsproducesdrivessupports another cycleBoard positioncurrent game stateMovechosen during self-playSelf-play gameexperienceGame outcomelearning signalEvaluator updateadjusted parameters
How does TD-Gammon generate experience, use outcomes to update its strategy, and repeat the process?

Following one learning cycle

Trace the role of self-play in a simplified TD-Gammon learning cycle.

1. Generate experience: TD-Gammon plays a game against itself. The resulting game supplies experience without requiring an expert to specify every move.

2. Evaluate positions: The program uses its nonlinear evaluator to assign predicted values to positions encountered during the game.

3. Compare prediction and later information: TD learning uses temporal information from the game to produce TD errors. These errors identify where the evaluator's predictions need adjustment.

4. Update the evaluator: The errors are backpropagated through the multilayer neural network, changing the network's learned evaluation.

5. Repeat: The updated program plays more self-play games, creating another cycle of experience and adjustment.

Self-play is the experience generator in the learning loop. It does not merely test a finished program; it supplies the games from which TD-Gammon learns.

Three Parts of the Learning Method

TD-Gammon's method is easiest to understand as a division of labor among three ideas. Self-play supplied experience. TD(λ) supplied the temporal-difference learning method. A multilayer neural network supplied nonlinear function approximation. TD errors connected the learning signal to the network, and backpropagation carried that signal through the network to update its evaluation.

ComponentRole in TD-Gammon
Self-playSupplied the games and positions used as learning experience.
TD(λ)Provided the temporal-difference learning method used to learn from game experience.
TD errorsProvided the training signal representing the adjustment needed by the evaluator.
Multilayer neural networkProvided nonlinear function approximation for evaluating positions.
BackpropagationTrained the network using the TD errors.

The learning roles in TD-Gammon's method

evaluatesleads to later informationcompared with targetdefines mismatchprocessed bytrainsBoard positiongame experiencePredicted valuenetwork outputTD errorprediction adjustmentNeural networknonlinear evaluatorUpdated targetlater game informationTD(λ)credit to earlier positions
How does a position's predicted value become a learning signal and then update a nonlinear evaluator?

Following the Learning Signal

What do you think happens?

Suppose the evaluator's prediction for a position does not agree with the later target produced from the game. What should happen next in the learning pipeline?

  • The mismatch is ignored because the game is already over.
  • The mismatch becomes a TD error used to train the evaluator.
  • The program replaces self-play with expert moves.
  • The position is removed from the learning process.
Reveal answer

Answer: The mismatch becomes a TD error used to train the evaluator.

TD-Gammon trained its multilayer neural network by backpropagating TD errors. The error therefore carries information from the game experience into an update of the nonlinear evaluator.

This simplified trace shows why the components belong together. A self-play position is evaluated by the neural network. Later information from the same game supplies an updated target. The difference between the prediction and that target forms a TD error. TD(λ) organizes the temporal-difference learning process, and backpropagation uses the resulting errors to train the nonlinear evaluator. The updated evaluator then participates in later self-play.

When explaining this method, describe the information path in order: self-play supplies experience; the network evaluates positions; later game information produces TD errors; TD(λ) supplies the temporal-difference learning method; backpropagation trains the network; and the updated evaluator returns to self-play.

What the Results Established

The results mattered because they tested whether strong play could emerge from this learning arrangement. After about 300,000 games against itself, TD-Gammon 0.0 played approximately as well as the best previous backgammon computer programs. This was significant because those earlier programs had relied heavily on backgammon knowledge, carefully designed features, or expert training examples.

approximately reached strength ofappeared to approach strength ofTD-Gammon 0.0about 300,000 self-playgamesBest computerprogramsapproximately matchedTD-Gammon 3.0later versionWorld's best humanplayersappeared to approach
How did TD-Gammon's results change the assessment of self-play and learned evaluation?

The later result was even more striking. TD-Gammon 3.0 appeared to be at or very near the playing strength of the world's best human players, based on games against world-class humans and analyses by backgammon grandmasters. The source presents the stronger claim cautiously: it notes that the program may already have been the world champion, rather than stating that conclusion as an absolute fact.

The influence also went beyond measurement. Human players changed some of their own play after studying TD-Gammon. In certain opening positions, the program used an approach different from the prevailing human convention, and the best human players later adopted the program's approach.

Common Interpretation Mistakes

  • Saying that TD-Gammon learned without any designed learning method.

    Self-play supplied experience, but TD-Gammon also used TD(λ), TD errors, backpropagation, and a multilayer neural network for nonlinear function approximation.

    Fix: Explain self-play as the source of experience and the other components as the machinery that learned from that experience.

  • Treating TD(λ) as the neural network.

    TD(λ) was the temporal-difference learning method, while the multilayer neural network provided nonlinear function approximation.

    Fix: Keep the algorithm and the function approximator conceptually separate.

  • Claiming that TD-Gammon used no prior knowledge of any kind.

    The source specifically describes essentially zero backgammon knowledge. It does not claim that the program had no programming structure or learning method.

    Fix: Use the precise claim: TD-Gammon used essentially zero built-in backgammon knowledge.

  • Treating the human-player result as an unqualified world-champion title.

    The source says TD-Gammon 3.0 appeared to approach the strength of the world's best human players and cautiously notes that it may already have been world champion.

    Fix: Preserve the source's cautious wording when interpreting the result.

  • Describing self-play as only an evaluation test.

    Self-play supplied the experience used for learning.

    Fix: Describe self-play as part of the training loop: games generated experience, and that experience drove evaluator updates.

Practice the Information Path

MEDIUM

Explain TD-Gammon's learning method in five connected statements. Your explanation should identify the source of experience, the learning algorithm, the training signal, the nonlinear function approximator, and the process used to train that approximator.

Hints
  • Begin with the games TD-Gammon played against itself.
  • Name TD(λ) as the temporal-difference learning method.
  • Identify TD errors as the signal used to adjust the evaluator.
  • Name the multilayer neural network as the nonlinear function approximator.
  • Finish by explaining that backpropagation trained the network using the TD errors.
MEDIUM

Compare TD-Gammon with an earlier program that relied on expert moves and specially crafted backgammon features. State one difference in the source of training experience and one difference in the source of playing strength.

Hints
  • For training experience, contrast expert-supplied examples with self-play.
  • For playing strength, contrast extensive built-in backgammon knowledge with learning from essentially zero backgammon knowledge.
  • Use Neurogammon as the source's example of a neural-network approach trained on expert moves and specially crafted features.

Key Takeaways

  1. TD-Gammon differed from earlier strong programs because it learned strong play through self-play with essentially zero built-in backgammon knowledge.
  2. Self-play supplied the games and positions that became the program's learning experience.
  3. TD(λ) provided the temporal-difference learning method, while TD errors supplied the adjustment signals.
  4. A multilayer neural network provided nonlinear function approximation, and backpropagation trained it using TD errors.
  5. TD-Gammon 0.0 approximately matched the best earlier computer programs after about 300,000 self-play games; TD-Gammon 3.0 appeared to approach the strength of the world's best human players.

Key Takeaways

  • TD-Gammon's major contribution was its route to strong play, not only its playing strength.
  • It replaced heavy dependence on expert backgammon knowledge and expert training examples with experience generated through self-play.
  • TD(λ), TD errors, backpropagation, and a multilayer neural network worked together as one learning method.
  • Its results against computer programs and human players showed the potential of self-play and learned nonlinear evaluation.