Concepts / Special Cases in Reinforcement Learning: Afterstates and Games

Special Cases in Reinforcement Learning: Afterstates and Games

An afterstate describes the task immediately after an action.

  • Programming

From Decision to Result

In reinforcement learning, choosing an action is not the end of the process. The action produces a new situation. An afterstate is the task viewed immediately after that action has produced its result. This idea changes which situation receives the main attention, but it does not create a new task.

Tracing an Afterstate

The sequence is straightforward: begin with the original state, select an action, and observe the state produced by that action. That produced state is the afterstate. The original state describes the situation before the decision; the afterstate describes the situation immediately after the decision has taken effect.

decisionproducesOriginal stateBefore the decisionActionSelected eventAfterstateState produced by theaction
What changes when an action is applied to the original state, and what does the resulting afterstate contain?

A simple state trace

Trace a generated task from its original state through an action to its afterstate.

Original state: Imagine a task situation in which a location has two available cars.

Action: The selected action moves one car away from that location.

Produced state: After the action, the location has one available car. This produced state is the afterstate.

The original state and the afterstate are different descriptions of the task: one is before the action, and the other is immediately after it.

The afterstate is not another name for the original state. It is the result of applying the action to the original state.

Reframing the Task

To reformulate a task around afterstates, trace the ordinary task description through the action. First identify the state before the decision. Next identify the action. Then describe the state that exists after that action. The post-action description becomes the main object of the reformulated task.

leads toproducesproducesOriginal stateBefore actionAfterstateProduced by actionActionSelected eventActionStill part of task
How does the sequence of states, actions, and value estimates change when the task is represented using afterstates?
ItemOrdinary descriptionAfterstate reformulation
Original stateThe situation before the decisionStill identified as the starting situation
ActionThe event selected from the original stateStill included because it produces the new situation
Resulting stateThe state produced by the actionBecomes the main focus of the description

Original State and Afterstate

The most useful test is to ask when each description is true. The original state is true before the action is selected or applied. The afterstate is true immediately after the action has produced its result. If two descriptions refer to these different points in the sequence, they should not be treated as interchangeable.

action selectedchanges situationOriginal stateSituation before actionAfterstateSituation after actionActionSelected event
Which parts of the situation belong to the original state, and which description belongs to the state changed by the action?

When analyzing a task, write the original state, action, and afterstate as three separate entries before trying to reformulate anything. This simple separation prevents the common mistake of treating the afterstate as merely another label for the original state.

Why Convergence Can Improve

The potential advantage of an afterstate reformulation is faster convergence. The reason must be evaluated for the particular task rather than assumed in every problem. By making the state produced by the action the focus, the reformulation may give the learning process a more useful description of what the action has brought about before the task continues.

decisionproducesmay supportOriginal stateBefore decisionActionProduces resultAfterstatePost-action focusFaster convergencePossible task-specificbenefit
How can representing the shared result of an action before the next external event reduce redundant learning and speed convergence?

Applying the Idea to Jack's Car Rental

Jack's Car Rental is the task used in the source exercise for applying the afterstate idea. The useful starting point is not a detailed calculation but a careful reformulation of the task's ordinary description.

  1. Identify the state before the decision.
  2. Identify the action selected from that state.
  3. Describe the state produced immediately after the action.
  4. Treat that post-action description as the afterstate.
  5. Compare the ordinary description with the afterstate-centered description.
  6. Ask whether the afterstate formulation could help convergence occur more quickly in this task.

This comparison keeps the task unchanged while changing its viewpoint. The action still matters, because it produces the afterstate. The analytical question is whether focusing on that produced situation gives a learning advantage for Jack's Car Rental.

Check Your Understanding

EASY

A task has an original state, a selected action, and a state produced by that action. Explain which of the three is the afterstate, how it differs from the original state, and why an afterstate reformulation does not remove the action from the task.

Hints
  • Look for the state that exists immediately after the action.
  • Describe the original state as the situation before the decision.
  • Remember that the action is what produces the afterstate.
MEDIUM

For Jack's Car Rental, outline the ordinary state-and-action description and then rewrite it around the state produced by the action. Finally, state why faster convergence is only a possible, task-specific advantage.

Hints
  • Use the sequence original state, action, produced state.
  • Call the produced state the afterstate.
  • Do not claim that afterstates always improve learning.

Key Takeaways

  1. An afterstate is the state of the task immediately after an action.
  2. To find an afterstate, separate the original state, the action, and the state produced by that action.
  3. An afterstate reformulation changes the focus of the task but does not remove actions.
  4. The original state describes the situation before the decision; the afterstate describes the situation after the action.
  5. Afterstate reformulation may speed convergence, but that benefit must be evaluated for the specific task.

Key Takeaways

  • An afterstate describes the task immediately after an action.
  • The original state, action, and afterstate are three distinct items in the sequence.
  • Reformulating around afterstates means making the action's resulting state the main focus.
  • Actions remain part of the task because they produce afterstates.
  • Faster convergence is a possible advantage that must be judged for the particular task.