Concepts / Watson's Daily-Double Wagering Strategy

Watson's Daily-Double Wagering Strategy

Watson evaluated wagers as actions rather than judging them only by their point amounts.

  • Programming

From Point Amount to Game Outcome

A Daily-Double wager can be judged in two different ways. One approach asks only how many points the wager might add. Watson used a broader question: if this particular legal wager is selected from the current game state, how favorable is the resulting chance of winning the game? That change in viewpoint turns wagering into an action-selection problem.

An action value is an estimate of the probability of winning from the current game state if a particular Daily-Double wager is selected.

Action Value as a Decision

choosecreatesevaluateCurrent statesLegal betround-dollar wagerAfterstateresulting game situationAction valueestimated chance of winning
What changes when Watson evaluates a Daily-Double wager as an action rather than as only a point amount?

Watson treated each legal wager as an action. For a candidate action, it estimated the value of the game from the current state if that action were taken. The resulting action value therefore refers to a possible decision, not merely to a number of points attached to the bet.

The phrase current game state matters: the same wager amount can have different consequences in different states, so Watson evaluated a wager in the context where it was available.

What the Afterstate Captures

To evaluate one candidate wager, begin with the current state s and select a legal round-dollar bet. That selection creates an afterstate: the game state resulting from choosing the action. The afterstate gives the evaluation a concrete situation to assess instead of leaving the wager as an isolated point amount.

choosecreatesestimateCurrent statebefore selecting a betAfterstateresulting game situationSelected betlegal round-dollar wagerState valueestimated value ofsituation
What does the game state look like immediately after a possible bet, and why does that afterstate matter?

The afterstate is not the final game outcome. It is the resulting game situation used during evaluation. The state value function estimates how favorable that situation is. The complete next game state also depends on which square is selected next, so the more precise description is an expected next afterstate value.

Two Estimates Work Together

contributesestimates likelihoodaffectsAfterstate valuelearned state valueestimateIn-categoryconfidencep DDDaily-Double responsecorrect or incorrectAction valueestimated probability of awin
How are the current state value and confidence of answering the Daily Double combined to estimate a possible wager's value?

Watson needed two kinds of information. First, its learned state value function, written as v-hat with parameters theta, estimated the probability of winning from a game state. During a betting decision, this function estimated the values of the afterstates associated with the legal wagers.

Second, Watson used in-category Daily-Double confidence, written p DD. This estimated the likelihood that Watson would respond correctly to the Daily-Double clue, which had not yet been revealed. The action-value evaluation therefore considered both the value of the resulting situations and the likelihood of a correct response.

State value answers, “How favorable is the resulting situation?” In-category confidence answers, “How likely is a correct response to this unrevealed clue?” The action value brings these considerations together for the chosen wager.

One Candidate Wager

Following one wager through the evaluation

Suppose Watson is in a current game state and is considering one legal round-dollar Daily-Double wager. How is that wager evaluated?

Start at the current state: The evaluation begins with the current game state s, not with the wager amount considered in isolation.

Select the candidate action: Watson treats the legal wager as an action that may be taken from state s.

Construct the afterstate: Selecting the wager creates the afterstate: the game situation resulting from that action.

Estimate the resulting value: The learned state value function estimates how favorable the resulting afterstate is, using the more precise idea of an expected next afterstate value because the next square also matters.

Account for response confidence: Because the Daily-Double clue has not yet been revealed, in-category Daily-Double confidence represents the estimated likelihood of a correct response.

Assign an action value: These estimates produce an action value for the candidate wager: an estimated probability of winning from the current state if that wager is selected.

The candidate is evaluated as a possible action with possible response outcomes, rather than only as a possible point increase.

This sequence explains why the afterstate and confidence estimate belong in the same evaluation. A bet changes the situation, and the unrevealed clue introduces uncertainty about the response. Watson's action value represents the estimated winning chance after taking both features into account.

Comparing the Legal Bets

comparecomparecompareLegal bet Aaction valueGreatest action valueselected wagerLegal bet Baction valueLegal bet Caction value
How did Watson evaluate each legal wager and select the one with the greatest estimated action value?
identifyconsiderevaluatecombineselect largestCurrent stateLegal betsround-dollar actionsCandidate afterstatesValue estimatesstate value and confidenceAction valuesone for each candidateSelected betgreatest action value
What sequence did Watson follow from generating legal bets to choosing one?

Watson repeated the evaluation for the legal bets available in the current state. Each candidate received an estimated action value. The basic selection rule was to choose the legal bet with the greatest action value.

A Generated Comparison

CandidateEvaluation focusDecision meaning
Wager AIts resulting afterstate and response confidence are evaluatedOne possible action value
Wager BIts resulting afterstate and response confidence are evaluatedA second possible action value
Wager CIts resulting afterstate and response confidence are evaluatedA third possible action value

Generated illustration of the repeated comparison process; these are not reported Watson wager values.

Choosing the best estimated action

Imagine three legal wagers are available. How should a learner describe Watson's comparison without reducing the decision to the largest point amount?

Evaluate Wager A: Consider the afterstate created by Wager A and use the learned state value estimate together with in-category Daily-Double confidence to obtain its action value.

Evaluate Wager B: Perform the same evaluation for Wager B. Its action value depends on its own resulting situation and the same kind of response-confidence estimate.

Evaluate Wager C: Perform the same evaluation for Wager C rather than assuming that its point amount determines its value.

Compare the results: Compare the three estimated action values, because each represents an estimated chance of winning from the current state under that candidate action.

Select the candidate: Choose the legal wager with the greatest action value unless a risk-abatement measure changes the basic choice.

The decision is an action-value comparison across legal bets, followed by selection of the largest estimate subject to possible risk abatement.

Mistakes in Reading the Strategy

  • Treating the action value as the number of points Watson expects to gain.

    The action value is an estimate of the probability of winning from the current state if the wager is selected.

    Fix: Describe the wager as an action whose effect on the chance of winning is being estimated.

  • Ignoring the afterstate.

    The state value function is used to estimate the values of the afterstates produced by candidate wagers.

    Fix: Follow each candidate bet forward to the afterstate it creates before discussing its estimated value.

  • Using only the learned state value function.

    The evaluation also uses in-category Daily-Double confidence to estimate the likelihood of a correct response.

    Fix: Mention both the value of the resulting situations and the confidence of responding correctly.

  • Assuming the largest wager must be selected.

    Watson's basic choice was the legal bet with the greatest estimated action value, with possible changes from risk-abatement measures.

    Fix: Compare the estimated action values for the legal bets.

Practice the Evaluation Pipeline

MEDIUM

Explain how Watson would evaluate two legal Daily-Double wagers from the same current game state. Your explanation should use the terms current state, afterstate, learned state value function, in-category Daily-Double confidence, action value, and greatest estimated action value.

Hints
  • Begin by identifying the legal wagers available from the current state.
  • Explain that each wager creates its own afterstate.
  • Describe how the state value function and response confidence contribute to each candidate's evaluation.
  • End by stating the basic selection rule and its risk-abatement qualification.

What do you think happens?

If one legal wager has a larger point amount but another has a greater estimated action value, which wager does the basic strategy select?

  • Always the larger point amount
  • The wager with the greatest estimated action value
  • Neither wager can be evaluated
  • The wager selected at random
Reveal answer

Answer: The wager with the greatest estimated action value.

Watson evaluated wagers as actions and used the basic choice rule of selecting the legal bet with the greatest action value, although risk-abatement measures could alter that choice.

The Decision in Brief

  1. An action value estimates the probability of winning from the current game state if a particular Daily-Double wager is selected.
  2. A legal wager creates an afterstate, which is the resulting game situation evaluated by the learned state value function.
  3. In-category Daily-Double confidence represents the estimated likelihood of responding correctly to the unrevealed clue.
  4. Watson computed an action value for each legal wager, compared those values, and basically selected the wager with the greatest estimate.
  5. Risk-abatement measures could modify the basic greatest-action-value choice.

Key Takeaways

  • Watson treated a Daily-Double wager as an action whose value concerns the chance of winning, not merely the number of points at stake.
  • The learned state value function evaluated the afterstates created by possible wagers.
  • In-category Daily-Double confidence accounted for the likelihood of a correct response to the unrevealed clue.
  • Watson compared the action values of legal wagers and basically selected the greatest one, subject to possible risk-abatement measures.