Concepts / State Value Functions and Action Values

State Value Functions and Action Values

Watson treated Daily-Double wagering as a decision under uncertainty rather than as a fixed betting rule.

  • Programming

A wager is a decision

Finding a Daily-Double clue was only the beginning of Watson's decision. Watson also had to choose how much to wager. That choice depended on the entire game situation and on Watson's uncertainty about whether it could answer the unrevealed clue correctly. Watson therefore treated Daily-Double wagering as a decision under uncertainty rather than as a fixed betting rule.

evaluatesinformsFixed betting rulesame ruleLearned decisionstate and action valuesGame statecurrent situationWagerselected legal betpDDcategory confidence
What changes when Daily-Double wagering is selected from state and action values instead of from one fixed betting rule?

From situation to wager

When a betting decision was required, Watson compared action values for the legal bets. An action value estimated the chance of winning from the current state if Watson made that particular bet. The wager with the highest action value was usually selected because it represented the option with the best estimated chance of winning among the available legal actions.

comparedcomparedcomparedSmall wageraction valueSelected wagerhighest estimated valueMedium wageraction valueLarge wageraction value
How does Watson compare possible wager amounts and why does it usually choose the bet with the highest estimated action value?

Choosing among three legal wagers

Imagine that three legal wagers are available. For this generated illustration, Watson's action-value estimates rank the small wager below the medium wager, and the medium wager below the large wager.

List the actions: The candidate actions are the small wager, the medium wager, and the large wager.

Evaluate the actions: Each candidate receives an action-value estimate representing the estimated chance of winning from the current state after choosing that wager.

Compare the estimates: The large wager has the highest estimate in this illustration, so it is the leading choice.

Select the action: Watson usually selected the legal wager with the highest action value.

In this generated illustration, the large wager is selected because it has the highest estimated action value. The numerical estimates are intentionally not specified because the source describes their meaning, not particular values.

Two estimates with different jobs

The action-value comparison depended on two learned estimates that should not be confused. One estimate described the afterstate that would result from making a particular bet. This is the broader game consequence: it connects a possible wager with an estimate of the resulting situation. The second estimate was in-category Daily-Double confidence, written as pDD. It estimated how likely Watson was to answer the unrevealed clue correctly in that category.

evaluatesestimatessupportssupportsState valuebroader game consequencesAfterstateresult of a betAction valuebet comparisonpDDanswer likelihoodClue successunrevealed category
What does the overall state value represent, what does in-category confidence represent, and how do they play different roles?
EstimateQuestion it addressesRole in wagering
State value functionWhat are the broader consequences of reaching a resulting game situation?Evaluates the afterstate produced by a possible bet
In-category Daily-Double confidence, pDDHow likely is Watson to answer the unrevealed clue correctly in this category?Supplies clue-answer information when action values are computed
Action valueHow promising is this particular legal wager from the current state?Combines the relevant estimates for comparing candidate bets

The decision process

offerscreatesevaluatedinformshighest selectedCurrent game stateDaily Double foundLegal wagerscandidate actionsAfterstatesone per wagerAction valuesestimated winning chancesSelected wagerhighest action valuepDDcategory confidence
How does Watson move from the current game state to a wager decision under uncertainty?

The process begins with the current game situation after Watson finds a Daily Double. Watson considers the legal wagers, estimates the resulting afterstate for each possible bet, and brings in pDD to represent its expected ability to answer the clue in that category. These estimates support the action values. Watson then compares the action values and usually chooses the wager with the highest estimate.

Learning through simulated games

Watson's state value function was supplied by reinforcement learning. It was learned with nonlinear TD(lambda) implemented through a multilayer neural network. The network received feature vectors designed for Jeopardy!, and its weights were trained by backpropagating temporal-difference errors during simulated games. The training connected a representation of the game state with an estimate of Watson's chance of winning from that state.

presentstested inproducesupdateslearnsGame statefeature vectorWager actionlegal betSimulated gameexperienceTD errortraining signalNetwork weightsupdated estimatesState value estimatechance of winning
How do simulated Jeopardy! games generate experience that updates Watson's estimates of different wagers?

The wagering strategy adapted the reinforcement-learning approach used in Tesauro's TD-Gammon system. Watson trained on millions of simulated games against models of human players. This gave the learning process repeated game experience from which to estimate the value of game situations and support wagering decisions.

Features that inform the strategy

represents stateestimateslearnssupportssupportsJeopardy! featuresfeature vectorNeural networktrained weightsState valueafterstate estimateWagering strategyaction-value comparisonHistorical accuracycategory performancepDDcategory confidence
How do Watson's game-state and confidence features flow into the learned wagering strategy?

Watson's features were designed for Jeopardy!, so the learned state representation was tied to the game it played. Separately, pDD was estimated from Watson's historical accuracy on clues in different categories over a large number of games. The state representation and the category-confidence estimate therefore contributed different information to the wagering strategy.

When explaining Watson's wager, identify both levels of reasoning: first, the possible game's resulting situation; second, Watson's expected ability to answer the unrevealed clue. Leaving out either estimate loses an important part of the action-value decision.

Mistakes to avoid

  • Treating Daily-Double wagering as a fixed rule

    The wagering decision depended on the entire game situation and was treated as a decision under uncertainty.

    Fix: Explain that Watson compared action values for legal bets in the current state.

  • Defining an action value as the value of the wager alone

    An action value estimated the chance of winning from the current state if Watson made that particular bet.

    Fix: Connect each action value to both the current state and the selected legal wager.

  • Using pDD as a synonym for state value

    pDD estimated the likelihood of answering the unrevealed clue correctly in its category, while the state value function evaluated broader game consequences.

    Fix: Keep the clue-answer estimate and the afterstate estimate separate.

  • Ignoring the afterstate

    The action values also drew on the estimate of the afterstate resulting from a particular bet.

    Fix: Mention the resulting game situation before describing the final action-value comparison.

  • Claiming that simulated games directly supplied a fixed wagering table

    The source describes neural-network training with feature vectors and temporal-difference errors to learn state-value estimates.

    Fix: Describe simulated games as experience used to train the network and update its estimates.

Check your reasoning

MEDIUM

A learner says: Watson chose a wager only because it was confident about the category. Correct the explanation in two parts: identify what pDD contributes, then identify what the state value estimate contributes.

Hints
  • pDD concerns the unrevealed clue and Watson's expected answer accuracy in its category.
  • The state value function concerns the broader game consequences of a possible bet and its resulting afterstate.
  • Both estimates support the comparison of action values for legal wagers.

What do you think happens?

Suppose one legal wager has a higher pDD-related expectation but produces a less favorable estimated afterstate than another wager. Which estimate should be ignored?

  • The state value estimate
  • pDD
  • Neither estimate
Reveal answer

Answer: Neither estimate

The source keeps the two estimates separate because both contribute to the action-value comparison: the state value evaluates broader game consequences, while pDD estimates the chance of answering the clue correctly.

The complete picture

  1. Watson treated Daily-Double wagering as a decision under uncertainty, not as a fixed betting rule.
  2. For each legal wager, an action value estimated the chance of winning from the current state if Watson selected that wager.
  3. The action-value comparison used an afterstate estimate for the broader game consequences and pDD for Watson's expected ability to answer the unrevealed clue in its category.
  4. The state value function was learned through nonlinear TD(lambda), a multilayer neural network, Jeopardy!-specific feature vectors, and temporal-difference training during simulated games.
  5. Watson usually selected the legal wager with the highest action value because it had the highest estimated chance of winning.

Key Takeaways

  • Watson used reinforcement learning to make Daily-Double wagering a state-dependent decision.
  • An action value represented the estimated chance of winning from the current state after taking a particular legal wager.
  • The state value function evaluated broader game consequences, while pDD estimated answer confidence for the unrevealed clue's category.
  • Millions of simulated games, Jeopardy!-specific feature vectors, neural-network training, and temporal-difference errors contributed to the learned strategy.