Concepts / Afterstate Value

Afterstate Value

Action value computation evaluates a candidate bet by its expected resulting value.

  • Programming

From Bet to Expected Value

An action value is an estimate of how valuable an action is before its outcome is known. In the Daily-Double setting described for Watson, the action is a candidate bet. The bet can lead to different resulting states because Watson may respond correctly or incorrectly. Action value computation combines the estimated values of those possible resulting states, weighted by Watson’s in-category Daily-Double confidence.

add betsubtract betcorrect responseincorrect responseweight by p_DDweight by 1 − p_DDweightweightCurrent score S_WS_W + betv-hat(S_W + bet)Action valueweighted outcome valueCandidate betbetS_W − betv-hat(S_W − bet)p_DDcorrect-response confidence1 − p_DDincorrect-response weight
How do the current score, candidate bet, in-category Daily-Double confidence, and estimated outcome values combine into one expected action value?

Tracing the Two Outcomes

An Illustrative Daily-Double Bet

Suppose the current score is 100, the candidate bet is 20, the in-category Daily-Double confidence is 0.75, the estimated value after a correct response is 140, and the estimated value after an incorrect response is 60. What is the illustrative action value?

Find the correct-response state: A correct response adds the bet to the current score: 100 + 20 = 120. The learned value function supplies the estimated value associated with the corresponding correct-response state, given here as 140.

Find the incorrect-response state: An incorrect response subtracts the bet from the current score: 100 − 20 = 80. The learned value function supplies the estimated value associated with the corresponding incorrect-response state, given here as 60.

Weight the estimates: The correct-response estimate receives weight 0.75. The incorrect-response estimate receives weight 1 − 0.75, which is 0.25.

Combine the weighted estimates: The action value is (0.75 × 140) + (0.25 × 60) = 105 + 15 = 120.

The illustrative action value is 120. The numbers are generated for practice; they are not reported values for Watson.

Action value = p_DD × v-hat(S_W + bet) + (1 − p_DD) × v-hat(S_W − bet)

The important point is that the bet amount alone does not determine the action value. The computation considers the current score, the score changes associated with a correct or incorrect response, the learned value of each resulting state, and the confidence used to weight the two possibilities.

The Learned Value Function

The value function, written v-hat, is a previously learned function that gives an estimated value for a game state. During action value computation, it is queried rather than replaced. For a candidate bet, it provides one estimate for the state associated with a correct response and another estimate for the state associated with an incorrect response. The action value then combines those estimates using p_DD and 1 − p_DD.

producesassess with v-hatcombine estimatessum weighted valuesCandidate betOutcome statescorrect and incorrectv-hat queriesestimated state valuesConfidence weightsp_DD and 1 − p_DDAction valueweighted combination
Where does the learned value function enter the action-value calculation?

Action Value and Resulting-State Value

A resulting-state value is an estimate for one state after a particular outcome. In the bet computation, v-hat(S_W + bet) is the estimate for the correct-response state, while v-hat(S_W − bet) is the estimate for the incorrect-response state. An action value is different: it evaluates the candidate bet before knowing which of those outcomes will occur by combining both estimates according to p_DD and 1 − p_DD.

evaluate actionassess resulting stateState-action pairbefore outcomeAction valueexpected resulting valueAfterstatestate after actionAfterstate valuevalue of one resultingstate
What is the difference between evaluating a state-action pair before an outcome is known and evaluating one state produced after the action?

Assessing the Afterstate

An afterstate value function assesses a state after an action has been taken. Its input is the resulting state, called the afterstate, rather than the original state-action pair. It therefore maps afterstates to value estimates.

This changes the unit being assessed. An ordinary action-value function may assess a state-action pair directly. An afterstate value function assesses the state produced by that pair. If several different state-action pairs produce the same afterstate, the function can use one shared value assessment for that resulting state.

producesinputmaps toState-action pairAfterstatestate after actionAfterstate valuefunctionValue estimateexpected future value
How does an afterstate value function map the state produced by an action to its expected future value?

One Afterstate, Many Routes

In reinforcement learning, different position-move pairs can sometimes lead to the same resulting position. The two pairs may begin from different positions or use different moves, yet the resulting position is identical. Because the resulting position is the same, both pairs have the same value assessment when value is assigned through the afterstate.

producesproducesassess oncePosition-move pairAShared afterpositionShared valueassessmentPosition-move pairB
How do multiple actions that lead to the same resulting state converge on one shared afterstate value assessment?

An ordinary action-value function may assess the two position-move pairs separately. An afterstate value function instead assesses their common resulting position directly. This avoids redundant work because the shared afterstate does not need separate value assessments merely because it was reached through different routes.

Transfer Through Shared Results

The efficiency benefit is also a learning benefit. When two different state-action pairs produce the same afterstate, learning about one pair transfers immediately to the other through that shared result. The afterstate is the link: both routes point to the same value assessment.

producesproducesupdate shared valueapplies to bothState-action pair AIdentical afterstateValue learningshared assessmentTransferred valueavailable to both routesState-action pair B
How can learning from two different state-action pairs update the same value when both pairs produce an identical afterstate?

Mistakes in Value Assessment

  • Treating the action value as the value of the bet amount alone.

    The action value also depends on the current score, the estimated values of the correct and incorrect outcome states, and the confidence used to weight those outcomes.

    Fix: Trace both resulting states and combine their estimated values using p_DD and 1 − p_DD.

  • Replacing the value function with the action value calculation.

    The value function estimates the values of game states. Action value computation queries it for outcome-state estimates and then combines them.

    Fix: Keep the sequence clear: produce outcome states, query v-hat, weight the estimates, and combine them.

  • Confusing one afterstate value with an action value.

    That is one resulting-state estimate. The action value accounts for both correct and incorrect response outcomes.

    Fix: Use both v-hat(S_W + bet) and v-hat(S_W − bet) with their respective weights.

  • Assessing identical afterstates separately.

    An afterstate value function assesses the shared resulting position directly, so separate assessments repeat the same work.

    Fix: Use the common afterstate as the shared input to the value function.

Check Your Understanding

EASY

A candidate bet has an in-category Daily-Double confidence of 0.60. The learned value function estimates the correct-response state at 90 and the incorrect-response state at 30. Compute the illustrative action value using the weighted-outcome structure.

Hints
  • The incorrect-response weight is 1 − 0.60.
  • Multiply each estimated state value by its corresponding weight.
  • Add the two weighted values.
MEDIUM

Two different position-move pairs produce one identical afterposition. Should an afterstate value function assign two separate value assessments or one shared assessment? Explain how learning from one pair can affect the other.

Hints
  • The afterstate value function takes the resulting state as its input.
  • Focus on whether the resulting afterposition is identical.
  • Relate the shared afterposition to redundant assessment and transfer.

Key Takeaways

  • An action value estimates the expected value of a candidate action before its outcome is known.
  • For a candidate bet, the computation uses the current score, the bet, the estimated values after correct and incorrect responses, and p_DD with 1 − p_DD as weights.
  • A resulting-state value assesses one state, while an action value combines the values of possible resulting states.
  • An afterstate value function maps the state produced after an action to a value estimate instead of assessing every state-action pair separately.
  • When different state-action pairs produce the same afterstate, they can share one assessment, allowing learning to transfer between them.