Self-Play
TD-Gammon combined self-play with TD learning to improve its backgammon performance.
The Learning Loop
Self-play means that a system generates experience by playing against another copy of itself. TD-Gammon used this approach together with temporal-difference learning, or TD learning, to improve its backgammon performance. The essential division of labor was simple: self-play supplied games and positions from which the program could learn, while TD learning supplied the learning method.
One Learning Cycle
From a Self-Played Game to an Updated Estimate
Trace what happens when TD-Gammon plays a backgammon game against another copy of itself.
1. Generate experience: The program plays a game against another copy of itself. This produces a sequence of backgammon positions rather than requiring a human opponent to provide every training game.
2. Observe positions and rewards: The self-play game supplies positions and game results that can be used as learning experience.
3. Apply TD learning: TD learning uses that experience to update the program's value estimates. Self-play is the source of experience; TD learning is the method that changes the estimates.
4. Repeat: The updated program plays more self-play games, creating another cycle of experience and learning.
Self-play and TD learning work together, but they are not the same thing: one generates experience and the other learns from it.
TD-Gammon's success did not mean that it relied only on manually supplied backgammon knowledge. TD-Gammon 0.0 began with zero backgammon knowledge, and its success motivated later additions of specialized features without abandoning self-play TD learning.
TD-Gammon's Progression
TD-Gammon developed through versions that added capabilities while preserving the self-play TD-learning approach. Version 0.0 was notable because it had zero backgammon knowledge. Later versions added specialized backgammon features, hidden units, and selective search. The progression therefore combined continued learning from self-play with increasingly specialized ways to represent positions or choose moves.
The result was more than an experimental improvement. Gerald Tesauro played the programs in a significant number of games against world-class human players. Based on those results and analyses by backgammon grandmasters, TD-Gammon 3.0 appeared to be at, or very near, the playing strength of the best human players in the world. The source even notes that it may already have been the world champion.
Selective Search
Selective search changed how TD-Gammon selected moves, not how it learned from self-play. In TD-Gammon 2.0 and 2.1, the program first considered candidate moves at one ply. It then carried out a second ply only for candidates ranked highly after that first ply. To save computer time, this deeper search was used for about four or five moves on average rather than for every candidate. TD-Gammon 3.0 extended the approach to selective three-ply search.
Relevant and Unusual Experience
Imperfect information exists when part of the situation that influences play is hidden or unknown to a contestant. In Jeopardy!, contestants cannot observe opponents' confidence about clues in different categories. Backgammon provides a contrasting setting in the source discussion: the board position is available to the players, so the issue is not the same hidden-confidence problem described for Jeopardy!.
Self-play can expose a system to states that are unusual in play against human contestants. This matters because a self-play opponent may behave differently from the human opponents the system will actually face. Watson was substantially different from human Jeopardy! contestants, so Watson-versus-Watson games would explore atypical regions of the state space rather than reliably reproducing the situations Watson would encounter against people.
Why Watson Needed Human-Relevant Experience
The key question is not whether self-play can generate experience. It can. The key question is whether that experience resembles the situations the system will face in the target game. For Watson, the answer was not sufficiently strong. Watson differed greatly from human contestants, and Jeopardy! included imperfect information because a contestant could not observe an opponent's confidence about clues in different categories.
Common Reasoning Errors
Treating self-play and TD learning as the same process.
Self-play generated the games and positions, while TD learning was the learning method applied to that experience.
Fix:
Describe self-play as the experience generator and TD learning as the method that learns from the experience.Assuming that selective search changed how TD-Gammon learned.
Selective search changed move selection. The learning process continued as before.
Fix:
Explain that the program searched more deeply for selected candidate moves while continuing to learn through self-play TD learning.Assuming that self-play is always representative of real opponents.
Watson differed substantially from human contestants, and contestants could not observe opponents' confidence about clues.
Fix:
Check whether the self-play opponent and its information viewpoint match the opponents and information conditions of the target game.Calling every hidden game detail imperfect information without identifying its effect on play.
The important issue is the difference between the full situation influencing play and the portion a contestant can observe.
Fix:
Define the hidden information precisely and explain why it affects the relevance of self-play experience.
Check Your Understanding
Explain in a short paragraph why self-play was highly successful for TD-Gammon but was not used to learn Watson's critical value function. Your answer should mention the source of experience, the learning method, opponent similarity, and imperfect information.
Hints
- Separate the role of self-play from the role of TD learning.
- Mention that TD-Gammon reached world-class backgammon strength through self-play TD learning.
- For Watson, focus on the difference between Watson and human contestants and the fact that opponent confidence was hidden.
| Question | TD-Gammon | Watson in Jeopardy! |
|---|---|---|
| What generated experience? | Self-play games | Self-play would generate Watson-versus-Watson games |
| How relevant was the opponent? | Self-play supported successful backgammon learning | Watson differed substantially from human contestants |
| What information issue mattered? | The source emphasizes board positions and backgammon play | Contestants could not observe opponents' confidence about clues |
| What was the consequence? | World-class playing strength and influence on expert strategy | Self-play was not used to learn the critical value function |
Comparing when self-play produced useful experience and when it did not
Key Takeaways
- TD-Gammon combined self-play, which generated games and positions, with TD learning, which learned from that experience.
- TD-Gammon progressed from version 0.0 with zero backgammon knowledge to later versions with specialized features, hidden units, and selective search.
- Selective search examined deeper move sequences only for highly ranked candidates; TD-Gammon 3.0 extended the approach to selective three-ply search.
- TD-Gammon reached world-class playing strength and influenced how expert human backgammon players approached certain positions.
- Self-play is useful only when its experience resembles the opponents, situations, and information conditions of the target game; this limitation prevented its use for learning Watson's critical value function.