State Value Functions and Action Values
Watson treated Daily-Double wagering as a decision under uncertainty rather than as a fixed betting rule.
A wager is a decision
Finding a Daily-Double clue was only the beginning of Watson's decision. Watson also had to choose how much to wager. That choice depended on the entire game situation and on Watson's uncertainty about whether it could answer the unrevealed clue correctly. Watson therefore treated Daily-Double wagering as a decision under uncertainty rather than as a fixed betting rule.
From situation to wager
When a betting decision was required, Watson compared action values for the legal bets. An action value estimated the chance of winning from the current state if Watson made that particular bet. The wager with the highest action value was usually selected because it represented the option with the best estimated chance of winning among the available legal actions.
Choosing among three legal wagers
Imagine that three legal wagers are available. For this generated illustration, Watson's action-value estimates rank the small wager below the medium wager, and the medium wager below the large wager.
List the actions: The candidate actions are the small wager, the medium wager, and the large wager.
Evaluate the actions: Each candidate receives an action-value estimate representing the estimated chance of winning from the current state after choosing that wager.
Compare the estimates: The large wager has the highest estimate in this illustration, so it is the leading choice.
Select the action: Watson usually selected the legal wager with the highest action value.
In this generated illustration, the large wager is selected because it has the highest estimated action value. The numerical estimates are intentionally not specified because the source describes their meaning, not particular values.
Two estimates with different jobs
The action-value comparison depended on two learned estimates that should not be confused. One estimate described the afterstate that would result from making a particular bet. This is the broader game consequence: it connects a possible wager with an estimate of the resulting situation. The second estimate was in-category Daily-Double confidence, written as pDD. It estimated how likely Watson was to answer the unrevealed clue correctly in that category.
| Estimate | Question it addresses | Role in wagering |
|---|---|---|
| State value function | What are the broader consequences of reaching a resulting game situation? | Evaluates the afterstate produced by a possible bet |
| In-category Daily-Double confidence, pDD | How likely is Watson to answer the unrevealed clue correctly in this category? | Supplies clue-answer information when action values are computed |
| Action value | How promising is this particular legal wager from the current state? | Combines the relevant estimates for comparing candidate bets |
The decision process
The process begins with the current game situation after Watson finds a Daily Double. Watson considers the legal wagers, estimates the resulting afterstate for each possible bet, and brings in pDD to represent its expected ability to answer the clue in that category. These estimates support the action values. Watson then compares the action values and usually chooses the wager with the highest estimate.
Learning through simulated games
Watson's state value function was supplied by reinforcement learning. It was learned with nonlinear TD(lambda) implemented through a multilayer neural network. The network received feature vectors designed for Jeopardy!, and its weights were trained by backpropagating temporal-difference errors during simulated games. The training connected a representation of the game state with an estimate of Watson's chance of winning from that state.
The wagering strategy adapted the reinforcement-learning approach used in Tesauro's TD-Gammon system. Watson trained on millions of simulated games against models of human players. This gave the learning process repeated game experience from which to estimate the value of game situations and support wagering decisions.
Features that inform the strategy
Watson's features were designed for Jeopardy!, so the learned state representation was tied to the game it played. Separately, pDD was estimated from Watson's historical accuracy on clues in different categories over a large number of games. The state representation and the category-confidence estimate therefore contributed different information to the wagering strategy.
When explaining Watson's wager, identify both levels of reasoning: first, the possible game's resulting situation; second, Watson's expected ability to answer the unrevealed clue. Leaving out either estimate loses an important part of the action-value decision.
Mistakes to avoid
Treating Daily-Double wagering as a fixed rule
The wagering decision depended on the entire game situation and was treated as a decision under uncertainty.
Fix:
Explain that Watson compared action values for legal bets in the current state.Defining an action value as the value of the wager alone
An action value estimated the chance of winning from the current state if Watson made that particular bet.
Fix:
Connect each action value to both the current state and the selected legal wager.Using pDD as a synonym for state value
pDD estimated the likelihood of answering the unrevealed clue correctly in its category, while the state value function evaluated broader game consequences.
Fix:
Keep the clue-answer estimate and the afterstate estimate separate.Ignoring the afterstate
The action values also drew on the estimate of the afterstate resulting from a particular bet.
Fix:
Mention the resulting game situation before describing the final action-value comparison.Claiming that simulated games directly supplied a fixed wagering table
The source describes neural-network training with feature vectors and temporal-difference errors to learn state-value estimates.
Fix:
Describe simulated games as experience used to train the network and update its estimates.
Check your reasoning
A learner says: Watson chose a wager only because it was confident about the category. Correct the explanation in two parts: identify what pDD contributes, then identify what the state value estimate contributes.
Hints
- pDD concerns the unrevealed clue and Watson's expected answer accuracy in its category.
- The state value function concerns the broader game consequences of a possible bet and its resulting afterstate.
- Both estimates support the comparison of action values for legal wagers.
What do you think happens?
Suppose one legal wager has a higher pDD-related expectation but produces a less favorable estimated afterstate than another wager. Which estimate should be ignored?
Reveal answer
Answer: Neither estimate
The source keeps the two estimates separate because both contribute to the action-value comparison: the state value evaluates broader game consequences, while pDD estimates the chance of answering the clue correctly.
The complete picture
- Watson treated Daily-Double wagering as a decision under uncertainty, not as a fixed betting rule.
- For each legal wager, an action value estimated the chance of winning from the current state if Watson selected that wager.
- The action-value comparison used an afterstate estimate for the broader game consequences and pDD for Watson's expected ability to answer the unrevealed clue in its category.
- The state value function was learned through nonlinear TD(lambda), a multilayer neural network, Jeopardy!-specific feature vectors, and temporal-difference training during simulated games.
- Watson usually selected the legal wager with the highest action value because it had the highest estimated chance of winning.
Key Takeaways
- Watson used reinforcement learning to make Daily-Double wagering a state-dependent decision.
- An action value represented the estimated chance of winning from the current state after taking a particular legal wager.
- The state value function evaluated broader game consequences, while pDD estimated answer confidence for the unrevealed clue's category.
- Millions of simulated games, Jeopardy!-specific feature vectors, neural-network training, and temporal-difference errors contributed to the learned strategy.