Concepts / Watson's Question Answering and Game-Playing Architecture

Watson's Question Answering and Game-Playing Architecture

Watson treated Daily-Double wagering as a decision under uncertainty rather than as a fixed betting rule.

  • Programming

A Wager Is More Than a Number

Finding a Daily-Double clue was only the beginning of Watson's decision. Watson also had to decide how much to wager. That choice depended on the entire game situation, not only on the clue. Instead of following one fixed betting rule, Watson treated wagering as a decision under uncertainty.

Watson compared the legal bets available in the current situation. Each bet received an action value representing an estimate of the chance of winning from that situation, and Watson usually selected the bet with the highest action value.

presentsevaluated ascomparedCurrent game stateboard and scoresLegal betscandidate wagersAction valuesestimated winning chancesSelected wagerhighest value, usually
How does Watson move from the current game situation to possible Daily-Double bets and then select a wager under uncertainty?

Comparing Candidate Bets

An action value belongs to a possible bet in a particular game situation. It estimates the chance of winning that would result from making that bet from the current state. The value is therefore not a permanent rating of the bet itself. The same wager can have a different action value when the surrounding game situation changes.

CandidateWhat Watson evaluatesDecision role
Legal betThe possible game situation after making that betOne candidate action
Another legal betA different possible game situation after making that betAnother candidate action
Highest action valueThe estimated winning chance associated with the candidateUsually selected by Watson
maps tomaps tomaps tocomparedcomparedcomparedBet Alegal wagerAction value Awinning estimateBet Blegal wagerAction value Bwinning estimateBet Clegal wagerAction value Cwinning estimateSelected wagerhighest action value
How does each possible bet map to an action value, and why does Watson usually choose the wager with the highest value?

Two Estimates Behind One Value

The action-value comparison used two learned estimates. First, Watson evaluated the afterstate: the game situation that would result from making a particular bet. Second, it used in-category Daily-Double confidence, written as pDD. This estimate represented how likely Watson was to answer the unrevealed clue correctly in that category.

The state value function evaluates the broader game consequences of a possible bet. It represents Watson's estimated chance of winning from a game state or resulting afterstate.

In-category Daily-Double confidence, or pDD, estimates Watson's likelihood of answering the unrevealed Daily-Double clue correctly in its category. It was estimated from Watson's historical accuracy on clues in different categories over a large number of games.

evaluatesdescribescontributescontributesState valuefunctionbroader game consequencesResulting afterstategame situation after betpDDclue-answer confidenceClue categoryhistorical categoryaccuracyAction valuecandidate wager estimate
What does the overall state value represent, what does in-category confidence represent, and how do they provide different information for wagering?

An Illustrative Bet Comparison

Choosing Between Three Legal Bets

Suppose Watson has found a Daily-Double clue and three legal bets are available. How should the decision be analyzed?

Inspect the current situation: Watson begins with the current game state, including the situation that would lead to each possible wager being considered.

Evaluate each afterstate: For each candidate bet, Watson considers the resulting afterstate and uses the state value function to estimate the broader game consequences.

Estimate clue confidence: Watson also considers pDD, its estimated likelihood of answering the unrevealed clue correctly in that category.

Compare action values: The two estimates support an action value for each legal bet. Watson compares those action values rather than applying one fixed wager rule.

Select the leading candidate: The candidate with the highest action value is usually selected because its estimated chance of winning is highest among the compared legal bets.

The wager is chosen by comparing predicted consequences and clue-answer confidence. The numerical values in this example are intentionally omitted because the source describes the decision process, not a particular game position.

informsdefinescontributescontributesFixed betting rulesame rule across situationsCurrent game statesituation-specificinformationpDDcategory confidenceAfterstate valuebroader game consequencesAction-value strategybet selected fromcomparison
How does a fixed betting rule differ from a strategy that considers the current board, confidence, possible outcomes, and expected rewards?

Learning Through Simulated Games

Reinforcement learning supplied Watson's state value function. Watson trained a multilayer neural network with nonlinear TD(lambda) on millions of simulated games against models of human players. During these games, temporal-difference errors were backpropagated to adjust the network's weights.

The training connected a representation of a game state with an estimate of Watson's chance of winning from that state. The learned function could then contribute to the evaluation of a possible Daily-Double wager by assessing the afterstate produced by that wager.

generates decisionleads tohelps produceupdatesevaluates future statesSimulated gamestatemodel of playWagercandidate decisionGame outcomelater resultTD errortraining signalNeural networklearned state value
How do simulated game states, wagers, outcomes, and rewards flow through repeated trials to produce a learned wagering strategy?

Features Designed for Jeopardy!

The neural network did not learn from an unspecified game representation. It received feature vectors designed for Jeopardy!. These features represented the game state used by the wagering model, while the question-answering system supplied information relevant to Watson's expected success on the unrevealed clue.

Watson-specific information therefore entered the decision in two related ways. The learned state value function represented the broader consequences of the game situation, and pDD represented Watson's category-specific likelihood of answering the clue correctly. Together, these estimates supported the action value for each possible legal wager.

processed byinformscontributescontributesJeopardy! featurevectorsgame representationState value functionafterstate estimateCategory accuracyhistorypast clue performancepDDcategory confidenceAction valuewager estimate
How do Watson's game and question-answering features enter the decision process and influence the predicted value of each possible wager?

Mistakes About Watson's Wager

  • Thinking that finding the Daily-Double clue determines the whole decision.

    Finding the clue was only the beginning. Watson still had to choose how much to wager.

    Fix: Separate clue identification from wager selection. The wager depends on the entire game situation.

  • Treating pDD as the complete value of a bet.

    pDD estimates the chance of answering the unrevealed clue correctly in its category. It does not by itself evaluate the broader game consequences.

    Fix: Use pDD as one learned estimate and keep it separate from the state value function.

  • Treating the state value function as a clue-answer confidence score.

    The state value function evaluates the broader game consequences of a state or afterstate.

    Fix: Interpret the state value as an estimate of winning from the game situation, not as category-specific clue confidence.

  • Assuming Watson used one fixed wager rule.

    Watson treated wagering as a decision under uncertainty and compared action values for legal bets.

    Fix: Analyze the current situation, the resulting afterstates, and pDD before comparing candidate wagers.

Check Your Reasoning

MEDIUM

Explain why an action value cannot be understood as only Watson's confidence in answering the Daily-Double clue. In your answer, identify the separate role of the state value function and the separate role of pDD.

Hints
  • Start with what the state value function evaluates.
  • Then describe what pDD estimates.
  • Finish by explaining why both estimates contribute to comparing legal bets.
MEDIUM

Trace the learning pipeline in order: simulated game state, wager, later outcome, temporal-difference error, neural-network update, and learned state value. Then explain how that learned value can later help evaluate a possible Daily-Double wager.

Hints
  • The network was trained during millions of simulated games against models of human players.
  • Temporal-difference errors were backpropagated to adjust the network's weights.
  • A possible wager creates an afterstate that can be evaluated by the learned function.

The Decision in One Pass

  1. Watson encountered a Daily-Double decision in a particular game situation.
  2. It considered the legal bets available in that situation.
  3. For each possible bet, it evaluated the resulting afterstate with a learned state value function.
  4. It estimated pDD, its likelihood of answering the unrevealed clue correctly in that category.
  5. It used these estimates to compare action values for the candidate wagers.
  6. It usually selected the legal bet with the highest action value.
  7. The state value function had been learned through nonlinear TD(lambda), a multilayer neural network, and millions of simulated games against models of human players.

Key Takeaways

  • Watson treated Daily-Double wagering as a decision under uncertainty rather than as a fixed betting rule.
  • An action value estimated the chance of winning associated with a particular legal bet from the current game situation.
  • The state value function evaluated broader game consequences, while pDD estimated the chance of answering the unrevealed clue correctly in its category.
  • Reinforcement learning trained Watson's state value function through a multilayer neural network and millions of simulated games against models of human players.
  • Watson-specific Jeopardy! feature vectors and category-accuracy information helped connect game-state evaluation with question-answering confidence.