States, Actions, and Rewards
Episodic interaction has natural boundaries and restarts its time numbering in each episode.
One Interaction, Two Temporal Shapes
A reinforcement learning interaction does not always unfold in the same way. Sometimes the agent and environment naturally reach a boundary, and a new interaction begins. Each such interaction is an episode. In other problems, the interaction continues without a natural division into separate episodes. These are continuing tasks.
The useful question is not simply whether an interaction lasts for a long or short time. The key question is whether it naturally separates into distinct episodes. If it does, time can restart at the beginning of each episode. If it does not, one continuing time sequence is enough.
Tracking Time Across Episodes
An episodic task may need two time labels in a fully explicit description. The state at time t in episode i can be written as S_t,i. The same pattern applies to the action and reward: A_t,i identifies the action at that time and episode, while R_t,i identifies the reward.
The time label t identifies a position within an episode. The episode label i identifies which episode contains that position. Without the episode label, the same time value could refer to the first step of many different episodes. A continuing task does not need this extra label because it has no sequence of separate episodes to distinguish.
Reading the Episode Boundary
Identifying the Required Labels
Suppose a task is described as a series of separate episodes. You want to refer to the state at time 2 in episode 4, the action at that point, and the resulting reward.
State: Use S_2,4 to identify time 2 within episode 4.
Action: Use A_2,4 because the action occurs at the same time and in the same episode.
Reward: Use R_2,4 because the reward is also associated with time 2 in episode 4.
Interpretation: The first index locates the step inside an episode, while the second index identifies the episode itself.
The explicit notation distinguishes matching time steps across different episodes.
If the discussion concerns only one episode, or if the same statement is intended for all episodes, the episode index can be omitted. The shorter forms S_t, A_t, and R_t are therefore not competing definitions. They are compact versions of the fully explicit notation.
The Absorbing-State Convention
Episodic and continuing tasks appear different because an episode ends after a finite number of rewards, while a continuing task can be represented by an indefinitely extended sequence. A single notation can handle both by treating the end of an episode as entry into a special absorbing state.
After termination, the process is represented as remaining in the absorbing state. The later transitions are self-transitions: the state stays absorbing, and the rewards are zero. This converts the finite episodic description into a continuing sequence without adding further nonzero rewards after termination.
The absorbing state is a notation convention for the end of an episode. It does not turn the original task into a new kind of task; it gives the ended task a continuing representation so that the same general return notation can be used.
Why Zero Rewards Preserve Return
Consider the reward sequence from the source discussion: +1, +1, +1, followed by 0, 0, 0, and so on. The first three rewards occur before termination. Every later reward comes from the absorbing state and is zero.
Extending the sequence adds only zero rewards after termination. Those added rewards contribute nothing to the return, so the return is unchanged. The original finite episode and its absorbing-state continuation therefore describe the same accumulated reward from the episode's meaningful interaction.
Comparing the Finite and Extended Descriptions
Compare an episode represented by +1, +1, +1 with the same episode extended after termination by 0, 0, 0, and then more zero rewards.
Finite description: The episode contains the three rewards received before termination.
Extended description: The same three rewards are followed by zero rewards generated while the process remains in the absorbing state.
Return comparison: The added rewards are all zero, so they do not alter the return.
The finite episode and its zero-reward absorbing continuation have the same return.
Mistakes with Temporal Notation
Treating every task as if it naturally breaks into episodes.
Continuing interaction has no natural breakdown into separate episodes, so the additional label is not needed to distinguish episodes.
Fix:
First determine whether the interaction naturally separates into episodes. Use an episode index when that distinction matters.Assuming that S_t always identifies one unique state occurrence.
In an episodic task, time numbering restarts in each episode. The same time index can occur in multiple episodes.
Fix:
Use S_t,i when the episode must be identified explicitly.Interpreting omitted episode indices as proof that episodes do not exist.
Dropping the episode index is a deliberate shorthand used when the episode number adds no useful information.
Fix:
Treat S_t as a compact form of S_t,i when the context already identifies the episode or applies uniformly to episodes.Adding arbitrary post-termination rewards to the absorbing continuation.
The unifying convention uses zero rewards after entry into the absorbing state.
Fix:
Represent later rewards from the absorbing state as zero.
Check Your Interpretation
A task is described as having separate episodes. At time 3, you need to refer to the state in episode 2. Which notation preserves both pieces of information, and what does each index identify?
Hints
- Use the explicit two-index state notation.
- The first index identifies position within an episode.
- The second index identifies the episode.
An ended episode has rewards +1, +1, +1. It is extended with zero rewards while the process remains in an absorbing state. Does the return change? Explain why.
Hints
- Identify the rewards added after termination.
- Ask whether those added rewards contribute anything to the return.
Key Takeaways
- An episodic task naturally separates into episodes, while a continuing task does not.
- Fully explicit episodic notation can use two labels: time within the episode and episode identity.
- The episode index may be omitted as deliberate shorthand when it adds no useful information.
- An absorbing state represents termination as self-transitions with zero rewards.
- Adding zero rewards after an episode ends does not change the return, allowing episodic and continuing descriptions to share a general notation.
Key Takeaways
- Episodic interaction has natural boundaries and restarts time numbering in each episode.
- Continuing interaction has no natural division into separate episodes.
- The episode index distinguishes otherwise matching time steps that belong to different episodes.
- An absorbing state turns episode termination into a continuing sequence of self-transitions with zero rewards.
- Because the added post-termination rewards are zero, extending an episode in this way leaves its return unchanged.