Concepts / Self-Play

Self-Play

TD-Gammon combined self-play with TD learning to improve its backgammon performance.

  • Programming

The Learning Loop

Self-play means that a system generates experience by playing against another copy of itself. TD-Gammon used this approach together with temporal-difference learning, or TD learning, to improve its backgammon performance. The essential division of labor was simple: self-play supplied games and positions from which the program could learn, while TD learning supplied the learning method.

producesare paired withhelps drivesupportsstarts another cycleSelf-play gameTwo copies of TD-GammonGame positionsStates encountered duringplayRewardsResults from the gameTD updateValue estimates changeImproved playThe next games use updatedestimates
How does TD-Gammon generate games against itself, use resulting positions and rewards to update value estimates, and repeat the cycle?

One Learning Cycle

From a Self-Played Game to an Updated Estimate

Trace what happens when TD-Gammon plays a backgammon game against another copy of itself.

1. Generate experience: The program plays a game against another copy of itself. This produces a sequence of backgammon positions rather than requiring a human opponent to provide every training game.

2. Observe positions and rewards: The self-play game supplies positions and game results that can be used as learning experience.

3. Apply TD learning: TD learning uses that experience to update the program's value estimates. Self-play is the source of experience; TD learning is the method that changes the estimates.

4. Repeat: The updated program plays more self-play games, creating another cycle of experience and learning.

Self-play and TD learning work together, but they are not the same thing: one generates experience and the other learns from it.

TD-Gammon's success did not mean that it relied only on manually supplied backgammon knowledge. TD-Gammon 0.0 began with zero backgammon knowledge, and its success motivated later additions of specialized features without abandoning self-play TD learning.

TD-Gammon's Progression

TD-Gammon developed through versions that added capabilities while preserving the self-play TD-learning approach. Version 0.0 was notable because it had zero backgammon knowledge. Later versions added specialized backgammon features, hidden units, and selective search. The progression therefore combined continued learning from self-play with increasingly specialized ways to represent positions or choose moves.

developed intoincludedcontributed to stronger playTD-Gammon 0.0Zero backgammon knowledgeLater versionsSpecialized features andhidden unitsSelective searchMore selective moveevaluationWorld-class strengthComparable to the besthuman players
What capabilities and playing-strength improvements were added as TD-Gammon progressed?

The result was more than an experimental improvement. Gerald Tesauro played the programs in a significant number of games against world-class human players. Based on those results and analyses by backgammon grandmasters, TD-Gammon 3.0 appeared to be at, or very near, the playing strength of the best human players in the world. The source even notes that it may already have been the world champion.

Selective Search

Selective search changed how TD-Gammon selected moves, not how it learned from self-play. In TD-Gammon 2.0 and 2.1, the program first considered candidate moves at one ply. It then carried out a second ply only for candidates ranked highly after that first ply. To save computer time, this deeper search was used for about four or five moves on average rather than for every candidate. TD-Gammon 3.0 extended the approach to selective three-ply search.

expandsrankssearches further in 2.0 and 2.1extended in 3.0Current positionCandidate movesFirst-ply rankingHighly ranked movesAbout four or five onaverageSecond plyTD-Gammon 2.0 and 2.1Third plyTD-Gammon 3.0
How does selective search expand only promising move sequences, and how does its role change across TD-Gammon 2.0, 2.1, and 3.0?

Relevant and Unusual Experience

Imperfect information exists when part of the situation that influences play is hidden or unknown to a contestant. In Jeopardy!, contestants cannot observe opponents' confidence about clues in different categories. Backgammon provides a contrasting setting in the source discussion: the board position is available to the players, so the issue is not the same hidden-confidence problem described for Jeopardy!.

informsdoes not revealBackgammon boardBoard position availableClue categoryVisible game informationBackgammon playPosition-based decisionsOpponent confidenceUnknown to the contestant
What information is available to a player in backgammon, and what information is hidden from a Jeopardy! contestant?

Self-play can expose a system to states that are unusual in play against human contestants. This matters because a self-play opponent may behave differently from the human opponents the system will actually face. Watson was substantially different from human Jeopardy! contestants, so Watson-versus-Watson games would explore atypical regions of the state space rather than reliably reproducing the situations Watson would encounter against people.

producecan produceHuman contestantsTarget opponentsSystem copySelf-play opponentHuman-play statesSituations encounteredagainst peopleAtypical statesDifferent opponent behavior
How can self-play expose a program to positions that are rarely or never encountered against human contestants?

Why Watson Needed Human-Relevant Experience

The key question is not whether self-play can generate experience. It can. The key question is whether that experience resembles the situations the system will face in the target game. For Watson, the answer was not sufficiently strong. Watson differed greatly from human contestants, and Jeopardy! included imperfect information because a contestant could not observe an opponent's confidence about clues in different categories.

differs fromare compared withcan producefurther limits relevancemade self-play unreliable forWatsonTarget systemOpponent mismatchWatson differs from humansAtypical experienceWatson-versus-Watson statesCritical valuefunctionNot learned from self-playHuman contestantsActual opponentsHidden confidenceContestant cannot observeit
How do opponent mismatch and hidden confidence make Watson-versus-Watson experience unreliable for learning the critical value function?

Common Reasoning Errors

  • Treating self-play and TD learning as the same process.

    Self-play generated the games and positions, while TD learning was the learning method applied to that experience.

    Fix: Describe self-play as the experience generator and TD learning as the method that learns from the experience.

  • Assuming that selective search changed how TD-Gammon learned.

    Selective search changed move selection. The learning process continued as before.

    Fix: Explain that the program searched more deeply for selected candidate moves while continuing to learn through self-play TD learning.

  • Assuming that self-play is always representative of real opponents.

    Watson differed substantially from human contestants, and contestants could not observe opponents' confidence about clues.

    Fix: Check whether the self-play opponent and its information viewpoint match the opponents and information conditions of the target game.

  • Calling every hidden game detail imperfect information without identifying its effect on play.

    The important issue is the difference between the full situation influencing play and the portion a contestant can observe.

    Fix: Define the hidden information precisely and explain why it affects the relevance of self-play experience.

Check Your Understanding

MEDIUM

Explain in a short paragraph why self-play was highly successful for TD-Gammon but was not used to learn Watson's critical value function. Your answer should mention the source of experience, the learning method, opponent similarity, and imperfect information.

Hints
  • Separate the role of self-play from the role of TD learning.
  • Mention that TD-Gammon reached world-class backgammon strength through self-play TD learning.
  • For Watson, focus on the difference between Watson and human contestants and the fact that opponent confidence was hidden.
QuestionTD-GammonWatson in Jeopardy!
What generated experience?Self-play gamesSelf-play would generate Watson-versus-Watson games
How relevant was the opponent?Self-play supported successful backgammon learningWatson differed substantially from human contestants
What information issue mattered?The source emphasizes board positions and backgammon playContestants could not observe opponents' confidence about clues
What was the consequence?World-class playing strength and influence on expert strategySelf-play was not used to learn the critical value function

Comparing when self-play produced useful experience and when it did not

Key Takeaways

  • TD-Gammon combined self-play, which generated games and positions, with TD learning, which learned from that experience.
  • TD-Gammon progressed from version 0.0 with zero backgammon knowledge to later versions with specialized features, hidden units, and selective search.
  • Selective search examined deeper move sequences only for highly ranked candidates; TD-Gammon 3.0 extended the approach to selective three-ply search.
  • TD-Gammon reached world-class playing strength and influenced how expert human backgammon players approached certain positions.
  • Self-play is useful only when its experience resembles the opponents, situations, and information conditions of the target game; this limitation prevented its use for learning Watson's critical value function.