Concepts / State and Action Representations

State and Action Representations

Afterstate value functions assess a state after an action has been taken.

  • Programming

One Result, Several Routes

When an action is taken, the important position for later assessment may be the position that now exists, rather than the particular route used to reach it. In reinforcement learning, different state-action pairs can sometimes lead to the same resulting position. An afterstate value function focuses on that resulting position.

chooseproduceschooseproducesPosition AstateMove AactionPosition CafterstatePosition BstateMove Baction
How can different starting states and actions lead to the same resulting afterstate?

The Afterstate as the Input

An afterstate value function assesses a state after an action has been taken. Its input is the resulting afterstate, rather than the original state-action pair.

An ordinary action-value representation may assess state-action pairs separately. For example, it may keep one assessment for taking Move A in Position A and another assessment for taking Move B in Position B. An afterstate representation changes the unit being assessed: it assesses the position produced after the action. If both routes produce the same afterstate, they use the same value assessment.

same resultsame resultPosition A + Move Aseparate pair assessmentPosition Cshared afterstateassessmentPosition B + Move Bseparate pair assessment
What changes when value is assigned to the state produced after an action rather than directly to the state-action pair?

Tracing Two State-Action Pairs

Two Routes to One Afterstate

Consider two state-action pairs: Position A with Move A, and Position B with Move B. Suppose both actions produce Position C.

Identify the routes: The first route is Position A followed by Move A. The second route is Position B followed by Move B.

Identify the result: Both routes lead to Position C. Position C is the afterstate for each pair.

Choose the assessment unit: An afterstate value function assesses Position C directly instead of keeping separate assessments for the two routes.

Apply the shared assessment: Because both pairs produce the same afterstate, the value assessment for Position C applies to both pairs.

The two state-action pairs share one afterstate value assessment rather than requiring separate assessments of the same resulting position.

The key comparison is not whether the starting positions have the same label. They can be different positions, and the actions can also be different. What matters for this representation is that the resulting position is the same. In tic-tac-toe, the source describes this possibility as different position-move pairs producing the same afterposition.

What do you think happens?

If Position A with Move A and Position B with Move B both produce Position C, how many afterstate positions need to be assessed for these two routes?

  • Zero
  • One
  • Two
  • One assessment for each starting position, even when the result is the same
Reveal answer

Answer: One

The afterstate value function assesses the shared resulting position, Position C. The two routes can therefore use the same value assessment.

Why Redundancy Disappears

If an ordinary action-value function assesses two state-action pairs separately, it may perform the same kind of assessment twice when both pairs lead to one identical afterstate. An afterstate value function assesses that shared afterstate directly. This avoids redundant work because the resulting position is evaluated once as the common object of interest.

same resultsame resultPair Aassessment 1Position Cone shared assessmentPair Bassessment 2
How does evaluating one shared afterstate avoid separately assessing every state-action pair that produces it?

Learning Transfer Through the Result

Suppose learning produces a better value assessment for Position C after observing the route from Position A with Move A. If Position B with Move B also produces Position C, the same improved assessment applies to that second pair. The transfer does not depend on the two starting positions or actions being identical. It depends on their shared afterstate.

produces and informsshared value appliesPosition A + Move Alearning experiencePosition Clearned valuePosition B + Move Breceives shared assessment
How does learning the value of one afterstate update the assessment of multiple state-action pairs?

When analyzing an afterstate value function, trace the action first and then compare the resulting positions. Do not stop after comparing the original state-action pairs. The shared result is the part of the representation that explains both efficiency and transfer.

Mistakes in Representation Choice

  • Treating an afterstate value function as if it assessed the original state-action pair directly.

    The afterstate value function maps from afterstates to value estimates rather than mapping separately from state-action pairs.

    Fix: Follow the action to its result and identify that resulting position as the input being assessed.

  • Assuming different starting positions always require different value assessments.

    Different position-move pairs can produce the same afterposition, so their shared result can have one value assessment.

    Fix: Compare the afterstates. If the resulting position is the same, recognize the shared assessment.

  • Explaining efficiency only as fewer stored assessments and missing transfer.

    The main efficiency benefit also includes immediate transfer: learning about one pair can inform another pair when both produce the same afterposition.

    Fix: Explain both parts: the shared result avoids redundant assessment, and learning attached to that result applies across the pairs that produce it.

Apply the Shared-Result Test

MEDIUM

Three state-action pairs are under consideration. Pair 1 produces Position R, Pair 2 produces Position S, and Pair 3 also produces Position R. Explain which pairs can share an afterstate value assessment and describe what happens when learning improves the assessment of Position R.

Hints
  • Identify the afterstate produced by each pair.
  • Group together pairs with the same resulting position.
  • Describe the effect of improving the shared assessment for Position R.

Practice Check

Pair 1 and Pair 3 both produce Position R, while Pair 2 produces Position S.

Group matching results: Pair 1 and Pair 3 belong together because their afterstate is Position R. Pair 2 belongs to a different group because its afterstate is Position S.

Count shared assessments: Position R needs one shared afterstate assessment for Pair 1 and Pair 3. Position S has its own assessment for Pair 2.

Trace learning: If learning improves the assessment of Position R through Pair 1, that improved assessment transfers to Pair 3 because Pair 3 produces the same afterstate.

Pair 1 and Pair 3 share the value assessment for Position R, and learning about one transfers to the other through that shared afterstate.

Summary

  1. An afterstate value function assesses the state produced after an action has been taken.
  2. Its input is the afterstate, not a separate original state-action pair.
  3. Different state-action pairs can produce the same afterstate and therefore share one value assessment.
  4. This shared assessment avoids redundant work.
  5. Learning about one pair transfers to another pair when both produce the same afterstate.

Key Takeaways

  • Afterstate value functions assess resulting states after actions.
  • Several different state-action pairs can lead to one shared afterstate.
  • Assessing the shared afterstate avoids redundant assessments.
  • A value learned through one route transfers to other routes that produce the same afterstate.