Concepts / Expected Next State Value

Expected Next State Value

Action value computation evaluates a candidate bet by its expected resulting value.

  • Programming

A Bet Has More Than One Possible Result

When Watson considers a Daily-Double wager, it does not evaluate the wager only by looking at the amount bet. It estimates the value that could follow from taking that action. A candidate bet can lead to different next states because Watson may respond correctly or incorrectly to the Daily-Double clue. The action value combines the estimated values of those possible outcomes.

An action value is the expected resulting value of taking a candidate bet, rather than the value of only one outcome.

considercorrectincorrectp_DD1 − p_DDCurrent stateCandidate betCorrect responsev-hat(S_W + bet)Action valueIncorrect responsev-hat(S_W − bet)
How do the estimated values of the correct-response and incorrect-response outcomes combine into one expected value for a candidate bet?

Tracing the Candidate Bet

Begin with Watson’s current game state, represented by s, and focus on one legal round-dollar bet, represented by bet. The current score is written S_W. If the response is correct, the score-related state is represented as S_W + bet. If the response is incorrect, it is represented as S_W − bet. These two expressions identify the two outcome states whose values must be estimated.

++−−queryqueryp_DD and 1 − p_DDweightedweightedCurrent scoreS_WCorrect stateS_W + betCorrect valuev-hat(S_W + bet)Action valueCandidate betbetIncorrect stateS_W − betIncorrect valuev-hat(S_W − bet)DD confidencep_DD
How does data move from the current score, chosen bet, and confidence in an in-category Daily Double to the computed action value?

The in-category Daily-Double confidence is written p_DD. It represents Watson’s confidence for responding correctly in that Daily Double. The computation uses p_DD for the correct-response estimate and 1 − p_DD for the incorrect-response estimate. The two weights account for the two response outcomes.

The Value Function’s Job

The value function, written v-hat, is a previously learned function that gives an estimated value for a game state.

Action value computation does not replace the value function. Instead, it queries the value function twice for the candidate bet: once for the state associated with a correct response and once for the state associated with an incorrect response. The resulting estimates are then combined using the in-category Daily-Double confidence.

Action value = p_DD × v-hat(S_W + bet) + (1 − p_DD) × v-hat(S_W − bet)

branchesqueryreturnsp_DD and 1 − p_DDaddCandidate betOutcome statesValue functionv-hatState estimatesWeighted estimatesAction value
Where does the value function fit in the process of evaluating a candidate bet, and what does it estimate for each possible next state?

A Numerical Walkthrough

The following numbers are generated for practice and are not values reported for Watson. They illustrate the structure of the computation. Suppose a candidate bet produces an estimated value of 80 for the correct-response state and an estimated value of 20 for the incorrect-response state. Suppose also that the in-category Daily-Double confidence is 0.75.

Combining Two Outcome Estimates

Given p_DD = 0.75, v-hat(S_W + bet) = 80, and v-hat(S_W − bet) = 20, compute the illustrative action value.

Find the incorrect-response weight: The incorrect-response weight is 1 − p_DD, so it is 1 − 0.75 = 0.25.

Weight the correct-response estimate: Multiply the correct-response confidence by its estimated value: 0.75 × 80 = 60.

Weight the incorrect-response estimate: Multiply the incorrect-response weight by its estimated value: 0.25 × 20 = 5.

Add the weighted estimates: The action value is 60 + 5 = 65.

The illustrative action value is 65.

The result is not the value of the correct-response state alone and not the value of the incorrect-response state alone. It is the expected resulting value of choosing the candidate bet when both outcomes are included with their corresponding weights.

Action Value Versus State Value

A state value describes one particular game state. For the candidate bet, the value function supplies one estimate for the correct-response state and another estimate for the incorrect-response state. An action value is different: it evaluates the candidate bet before the response outcome is known by combining those possible state values according to p_DD and 1 − p_DD.

correct outcomeincorrect outcomeweightedweightedAction valueCandidate betCorrect state valuev-hat(S_W + bet)Incorrect state valuev-hat(S_W − bet)
What is the difference between valuing the candidate action before the outcome is known and valuing one specific state after a correct or incorrect response?
QuantityWhat it evaluatesHow it is obtained
State valueOne resulting game stateThe value function estimates that state
Action valueA candidate bet before its response outcome is knownCorrect- and incorrect-response state estimates are weighted and added

Mistakes in the Computation

  • Treating the action value as the wager amount

    The computation evaluates the expected resulting value of the wager, not only the amount wagered.

    Fix: Use the candidate bet to identify the possible outcome states, query the value function for both, and combine the estimates.

  • Using only the correct-response estimate

    The candidate bet can produce both correct and incorrect response outcomes.

    Fix: Include the incorrect-response estimate and weight the two estimates by p_DD and 1 − p_DD.

  • Using p_DD for both outcomes

    The incorrect-response outcome uses the complementary weight 1 − p_DD.

    Fix: Use p_DD for the correct-response estimate and 1 − p_DD for the incorrect-response estimate.

  • Confusing a state estimate with the action value

    That expression estimates only the state associated with a correct response.

    Fix: Reserve the action value for the weighted combination of the relevant outcome-state estimates.

When tracing a candidate bet, write the two branches explicitly before calculating: correct response leads to S_W + bet, and incorrect response leads to S_W − bet. Then attach the value-function estimate and the appropriate probability weight to each branch.

Check Your Understanding

MEDIUM

A generated practice case has p_DD = 0.60, a correct-response state estimate of 90, and an incorrect-response state estimate of 30. Compute the action value. Then state whether your result is a single state value or a weighted evaluation of the candidate bet.

Hints
  • First compute 1 − p_DD.
  • Multiply each state estimate by its corresponding weight.
  • Add the two weighted contributions.

What do you think happens?

Using p_DD = 0.60, a correct-response estimate of 90, and an incorrect-response estimate of 30, what is the action value?

  • 54
  • 66
  • 72
  • 90
Reveal answer

Answer: 66

The correct-response contribution is 0.60 × 90 = 54. The incorrect-response weight is 0.40, so its contribution is 0.40 × 30 = 12. Adding them gives an action value of 66.

The Evaluation Pattern

To evaluate a legal round-dollar bet, start with the current score and candidate bet, identify the correct- and incorrect-response states, ask the previously learned value function for an estimate of each state, and combine the estimates using the in-category Daily-Double confidence. This produces an action value: an expected evaluation of the candidate bet before the response outcome is known.

  1. Start with the current state and one legal round-dollar bet.
  2. Represent the correct-response state with S_W + bet.
  3. Represent the incorrect-response state with S_W − bet.
  4. Estimate both states with the value function v-hat.
  5. Weight the estimates by p_DD and 1 − p_DD.
  6. Add the weighted estimates to obtain the action value.

Key Takeaways

  • An action value evaluates a candidate bet by its expected resulting value.
  • The value function estimates the values of the correct-response and incorrect-response states.
  • The current score and candidate bet identify the two score-related outcome states, S_W + bet and S_W − bet.
  • The correct-response estimate is weighted by p_DD, while the incorrect-response estimate is weighted by 1 − p_DD.
  • An action value is a weighted combination of possible state values, not the value of one resulting state alone.