Expected Next State Value
Action value computation evaluates a candidate bet by its expected resulting value.
A Bet Has More Than One Possible Result
When Watson considers a Daily-Double wager, it does not evaluate the wager only by looking at the amount bet. It estimates the value that could follow from taking that action. A candidate bet can lead to different next states because Watson may respond correctly or incorrectly to the Daily-Double clue. The action value combines the estimated values of those possible outcomes.
An action value is the expected resulting value of taking a candidate bet, rather than the value of only one outcome.
Tracing the Candidate Bet
Begin with Watson’s current game state, represented by s, and focus on one legal round-dollar bet, represented by bet. The current score is written S_W. If the response is correct, the score-related state is represented as S_W + bet. If the response is incorrect, it is represented as S_W − bet. These two expressions identify the two outcome states whose values must be estimated.
The in-category Daily-Double confidence is written p_DD. It represents Watson’s confidence for responding correctly in that Daily Double. The computation uses p_DD for the correct-response estimate and 1 − p_DD for the incorrect-response estimate. The two weights account for the two response outcomes.
The Value Function’s Job
The value function, written v-hat, is a previously learned function that gives an estimated value for a game state.
Action value computation does not replace the value function. Instead, it queries the value function twice for the candidate bet: once for the state associated with a correct response and once for the state associated with an incorrect response. The resulting estimates are then combined using the in-category Daily-Double confidence.
Action value = p_DD × v-hat(S_W + bet) + (1 − p_DD) × v-hat(S_W − bet)
A Numerical Walkthrough
The following numbers are generated for practice and are not values reported for Watson. They illustrate the structure of the computation. Suppose a candidate bet produces an estimated value of 80 for the correct-response state and an estimated value of 20 for the incorrect-response state. Suppose also that the in-category Daily-Double confidence is 0.75.
Combining Two Outcome Estimates
Given p_DD = 0.75, v-hat(S_W + bet) = 80, and v-hat(S_W − bet) = 20, compute the illustrative action value.
Find the incorrect-response weight: The incorrect-response weight is 1 − p_DD, so it is 1 − 0.75 = 0.25.
Weight the correct-response estimate: Multiply the correct-response confidence by its estimated value: 0.75 × 80 = 60.
Weight the incorrect-response estimate: Multiply the incorrect-response weight by its estimated value: 0.25 × 20 = 5.
Add the weighted estimates: The action value is 60 + 5 = 65.
The illustrative action value is 65.
The result is not the value of the correct-response state alone and not the value of the incorrect-response state alone. It is the expected resulting value of choosing the candidate bet when both outcomes are included with their corresponding weights.
Action Value Versus State Value
A state value describes one particular game state. For the candidate bet, the value function supplies one estimate for the correct-response state and another estimate for the incorrect-response state. An action value is different: it evaluates the candidate bet before the response outcome is known by combining those possible state values according to p_DD and 1 − p_DD.
| Quantity | What it evaluates | How it is obtained |
|---|---|---|
| State value | One resulting game state | The value function estimates that state |
| Action value | A candidate bet before its response outcome is known | Correct- and incorrect-response state estimates are weighted and added |
Mistakes in the Computation
Treating the action value as the wager amount
The computation evaluates the expected resulting value of the wager, not only the amount wagered.
Fix:
Use the candidate bet to identify the possible outcome states, query the value function for both, and combine the estimates.Using only the correct-response estimate
The candidate bet can produce both correct and incorrect response outcomes.
Fix:
Include the incorrect-response estimate and weight the two estimates by p_DD and 1 − p_DD.Using p_DD for both outcomes
The incorrect-response outcome uses the complementary weight 1 − p_DD.
Fix:
Use p_DD for the correct-response estimate and 1 − p_DD for the incorrect-response estimate.Confusing a state estimate with the action value
That expression estimates only the state associated with a correct response.
Fix:
Reserve the action value for the weighted combination of the relevant outcome-state estimates.
When tracing a candidate bet, write the two branches explicitly before calculating: correct response leads to S_W + bet, and incorrect response leads to S_W − bet. Then attach the value-function estimate and the appropriate probability weight to each branch.
Check Your Understanding
A generated practice case has p_DD = 0.60, a correct-response state estimate of 90, and an incorrect-response state estimate of 30. Compute the action value. Then state whether your result is a single state value or a weighted evaluation of the candidate bet.
Hints
- First compute 1 − p_DD.
- Multiply each state estimate by its corresponding weight.
- Add the two weighted contributions.
What do you think happens?
Using p_DD = 0.60, a correct-response estimate of 90, and an incorrect-response estimate of 30, what is the action value?
Reveal answer
Answer: 66
The correct-response contribution is 0.60 × 90 = 54. The incorrect-response weight is 0.40, so its contribution is 0.40 × 30 = 12. Adding them gives an action value of 66.
The Evaluation Pattern
To evaluate a legal round-dollar bet, start with the current score and candidate bet, identify the correct- and incorrect-response states, ask the previously learned value function for an estimate of each state, and combine the estimates using the in-category Daily-Double confidence. This produces an action value: an expected evaluation of the candidate bet before the response outcome is known.
- Start with the current state and one legal round-dollar bet.
- Represent the correct-response state with S_W + bet.
- Represent the incorrect-response state with S_W − bet.
- Estimate both states with the value function v-hat.
- Weight the estimates by p_DD and 1 − p_DD.
- Add the weighted estimates to obtain the action value.
Key Takeaways
- An action value evaluates a candidate bet by its expected resulting value.
- The value function estimates the values of the correct-response and incorrect-response states.
- The current score and candidate bet identify the two score-related outcome states, S_W + bet and S_W − bet.
- The correct-response estimate is weighted by p_DD, while the incorrect-response estimate is weighted by 1 − p_DD.
- An action value is a weighted combination of possible state values, not the value of one resulting state alone.