State and Action Representations
Afterstate value functions assess a state after an action has been taken.
One Result, Several Routes
When an action is taken, the important position for later assessment may be the position that now exists, rather than the particular route used to reach it. In reinforcement learning, different state-action pairs can sometimes lead to the same resulting position. An afterstate value function focuses on that resulting position.
The Afterstate as the Input
An afterstate value function assesses a state after an action has been taken. Its input is the resulting afterstate, rather than the original state-action pair.
An ordinary action-value representation may assess state-action pairs separately. For example, it may keep one assessment for taking Move A in Position A and another assessment for taking Move B in Position B. An afterstate representation changes the unit being assessed: it assesses the position produced after the action. If both routes produce the same afterstate, they use the same value assessment.
Tracing Two State-Action Pairs
Two Routes to One Afterstate
Consider two state-action pairs: Position A with Move A, and Position B with Move B. Suppose both actions produce Position C.
Identify the routes: The first route is Position A followed by Move A. The second route is Position B followed by Move B.
Identify the result: Both routes lead to Position C. Position C is the afterstate for each pair.
Choose the assessment unit: An afterstate value function assesses Position C directly instead of keeping separate assessments for the two routes.
Apply the shared assessment: Because both pairs produce the same afterstate, the value assessment for Position C applies to both pairs.
The two state-action pairs share one afterstate value assessment rather than requiring separate assessments of the same resulting position.
The key comparison is not whether the starting positions have the same label. They can be different positions, and the actions can also be different. What matters for this representation is that the resulting position is the same. In tic-tac-toe, the source describes this possibility as different position-move pairs producing the same afterposition.
What do you think happens?
If Position A with Move A and Position B with Move B both produce Position C, how many afterstate positions need to be assessed for these two routes?
Reveal answer
Answer: One
The afterstate value function assesses the shared resulting position, Position C. The two routes can therefore use the same value assessment.
Why Redundancy Disappears
If an ordinary action-value function assesses two state-action pairs separately, it may perform the same kind of assessment twice when both pairs lead to one identical afterstate. An afterstate value function assesses that shared afterstate directly. This avoids redundant work because the resulting position is evaluated once as the common object of interest.
Learning Transfer Through the Result
Suppose learning produces a better value assessment for Position C after observing the route from Position A with Move A. If Position B with Move B also produces Position C, the same improved assessment applies to that second pair. The transfer does not depend on the two starting positions or actions being identical. It depends on their shared afterstate.
When analyzing an afterstate value function, trace the action first and then compare the resulting positions. Do not stop after comparing the original state-action pairs. The shared result is the part of the representation that explains both efficiency and transfer.
Mistakes in Representation Choice
Treating an afterstate value function as if it assessed the original state-action pair directly.
The afterstate value function maps from afterstates to value estimates rather than mapping separately from state-action pairs.
Fix:
Follow the action to its result and identify that resulting position as the input being assessed.Assuming different starting positions always require different value assessments.
Different position-move pairs can produce the same afterposition, so their shared result can have one value assessment.
Fix:
Compare the afterstates. If the resulting position is the same, recognize the shared assessment.Explaining efficiency only as fewer stored assessments and missing transfer.
The main efficiency benefit also includes immediate transfer: learning about one pair can inform another pair when both produce the same afterposition.
Fix:
Explain both parts: the shared result avoids redundant assessment, and learning attached to that result applies across the pairs that produce it.
Apply the Shared-Result Test
Three state-action pairs are under consideration. Pair 1 produces Position R, Pair 2 produces Position S, and Pair 3 also produces Position R. Explain which pairs can share an afterstate value assessment and describe what happens when learning improves the assessment of Position R.
Hints
- Identify the afterstate produced by each pair.
- Group together pairs with the same resulting position.
- Describe the effect of improving the shared assessment for Position R.
Practice Check
Pair 1 and Pair 3 both produce Position R, while Pair 2 produces Position S.
Group matching results: Pair 1 and Pair 3 belong together because their afterstate is Position R. Pair 2 belongs to a different group because its afterstate is Position S.
Count shared assessments: Position R needs one shared afterstate assessment for Pair 1 and Pair 3. Position S has its own assessment for Pair 2.
Trace learning: If learning improves the assessment of Position R through Pair 1, that improved assessment transfers to Pair 3 because Pair 3 produces the same afterstate.
Pair 1 and Pair 3 share the value assessment for Position R, and learning about one transfers to the other through that shared afterstate.
Summary
- An afterstate value function assesses the state produced after an action has been taken.
- Its input is the afterstate, not a separate original state-action pair.
- Different state-action pairs can produce the same afterstate and therefore share one value assessment.
- This shared assessment avoids redundant work.
- Learning about one pair transfers to another pair when both produce the same afterstate.
Key Takeaways
- Afterstate value functions assess resulting states after actions.
- Several different state-action pairs can lead to one shared afterstate.
- Assessing the shared afterstate avoids redundant assessments.
- A value learned through one route transfers to other routes that produce the same afterstate.