Concepts / States, Actions, and Rewards

States, Actions, and Rewards

Episodic interaction has natural boundaries and restarts its time numbering in each episode.

  • Programming

One Interaction, Two Temporal Shapes

A reinforcement learning interaction does not always unfold in the same way. Sometimes the agent and environment naturally reach a boundary, and a new interaction begins. Each such interaction is an episode. In other problems, the interaction continues without a natural division into separate episodes. These are continuing tasks.

The useful question is not simply whether an interaction lasts for a long or short time. The key question is whether it naturally separates into distinct episodes. If it does, time can restart at the beginning of each episode. If it does not, one continuing time sequence is enough.

separates intothen restartscontinues asEpisodic taskEpisode 1Ongoing sequenceEpisode 2Continuing task
How can you tell whether an interaction naturally ends and restarts as episodes or continues indefinitely?

Tracking Time Across Episodes

An episodic task may need two time labels in a fully explicit description. The state at time t in episode i can be written as S_t,i. The same pattern applies to the action and reward: A_t,i identifies the action at that time and episode, while R_t,i identifies the reward.

The time label t identifies a position within an episode. The episode label i identifies which episode contains that position. Without the episode label, the same time value could refer to the first step of many different episodes. A continuing task does not need this extra label because it has no sequence of separate episodes to distinguish.

next time stepnext time stepS_0,1episode 1, time 0S_1,1episode 1, time 1S_0,2episode 2, time 0S_1,2episode 2, time 1
How do time steps map to separate episodes when the time index restarts at the beginning of each episode?

Reading the Episode Boundary

Identifying the Required Labels

Suppose a task is described as a series of separate episodes. You want to refer to the state at time 2 in episode 4, the action at that point, and the resulting reward.

State: Use S_2,4 to identify time 2 within episode 4.

Action: Use A_2,4 because the action occurs at the same time and in the same episode.

Reward: Use R_2,4 because the reward is also associated with time 2 in episode 4.

Interpretation: The first index locates the step inside an episode, while the second index identifies the episode itself.

The explicit notation distinguishes matching time steps across different episodes.

If the discussion concerns only one episode, or if the same statement is intended for all episodes, the episode index can be omitted. The shorter forms S_t, A_t, and R_t are therefore not competing definitions. They are compact versions of the fully explicit notation.

The Absorbing-State Convention

Episodic and continuing tasks appear different because an episode ends after a finite number of rewards, while a continuing task can be represented by an indefinitely extended sequence. A single notation can handle both by treating the end of an episode as entry into a special absorbing state.

After termination, the process is represented as remaining in the absorbing state. The later transitions are self-transitions: the state stays absorbing, and the rewards are zero. This converts the finite episodic description into a continuing sequence without adding further nonzero rewards after termination.

terminationself-transitionrewardActive statebefore terminationAbsorbing stateafter termination0later reward
What happens to the state, actions, and rewards after an episode ends, and how does an absorbing state represent that continuation?

The absorbing state is a notation convention for the end of an episode. It does not turn the original task into a new kind of task; it gives the ended task a continuing representation so that the same general return notation can be used.

Why Zero Rewards Preserve Return

Consider the reward sequence from the source discussion: +1, +1, +1, followed by 0, 0, 0, and so on. The first three rewards occur before termination. Every later reward comes from the absorbing state and is zero.

contributes tocontributes to+1, +1, +1episode rewardsEpisode return+1, +1, +1, 0, 0, 0extended sequenceSame return
Why does adding zero-reward steps after an episode ends leave the total return unchanged?

Extending the sequence adds only zero rewards after termination. Those added rewards contribute nothing to the return, so the return is unchanged. The original finite episode and its absorbing-state continuation therefore describe the same accumulated reward from the episode's meaningful interaction.

Comparing the Finite and Extended Descriptions

Compare an episode represented by +1, +1, +1 with the same episode extended after termination by 0, 0, 0, and then more zero rewards.

Finite description: The episode contains the three rewards received before termination.

Extended description: The same three rewards are followed by zero rewards generated while the process remains in the absorbing state.

Return comparison: The added rewards are all zero, so they do not alter the return.

The finite episode and its zero-reward absorbing continuation have the same return.

Mistakes with Temporal Notation

  • Treating every task as if it naturally breaks into episodes.

    Continuing interaction has no natural breakdown into separate episodes, so the additional label is not needed to distinguish episodes.

    Fix: First determine whether the interaction naturally separates into episodes. Use an episode index when that distinction matters.

  • Assuming that S_t always identifies one unique state occurrence.

    In an episodic task, time numbering restarts in each episode. The same time index can occur in multiple episodes.

    Fix: Use S_t,i when the episode must be identified explicitly.

  • Interpreting omitted episode indices as proof that episodes do not exist.

    Dropping the episode index is a deliberate shorthand used when the episode number adds no useful information.

    Fix: Treat S_t as a compact form of S_t,i when the context already identifies the episode or applies uniformly to episodes.

  • Adding arbitrary post-termination rewards to the absorbing continuation.

    The unifying convention uses zero rewards after entry into the absorbing state.

    Fix: Represent later rewards from the absorbing state as zero.

Check Your Interpretation

EASY

A task is described as having separate episodes. At time 3, you need to refer to the state in episode 2. Which notation preserves both pieces of information, and what does each index identify?

Hints
  • Use the explicit two-index state notation.
  • The first index identifies position within an episode.
  • The second index identifies the episode.
MEDIUM

An ended episode has rewards +1, +1, +1. It is extended with zero rewards while the process remains in an absorbing state. Does the return change? Explain why.

Hints
  • Identify the rewards added after termination.
  • Ask whether those added rewards contribute anything to the return.

Key Takeaways

  1. An episodic task naturally separates into episodes, while a continuing task does not.
  2. Fully explicit episodic notation can use two labels: time within the episode and episode identity.
  3. The episode index may be omitted as deliberate shorthand when it adds no useful information.
  4. An absorbing state represents termination as self-transitions with zero rewards.
  5. Adding zero rewards after an episode ends does not change the return, allowing episodic and continuing descriptions to share a general notation.

Key Takeaways

  • Episodic interaction has natural boundaries and restarts time numbering in each episode.
  • Continuing interaction has no natural division into separate episodes.
  • The episode index distinguishes otherwise matching time steps that belong to different episodes.
  • An absorbing state turns episode termination into a continuing sequence of self-transitions with zero rewards.
  • Because the added post-termination rewards are zero, extending an episode in this way leaves its return unchanged.