Episodic and Continuing Tasks
Episodic tasks have interactions that naturally divide into separate episodes.
One Run or an Ongoing Interaction
A reinforcement-learning interaction can have a clear ending, or it can continue without a natural point at which one interaction ends and another begins. This distinction produces two task types: episodic tasks and continuing tasks. The key question is whether the interaction naturally divides into complete runs with identifiable beginnings and endings.
An episodic task is organized into separate episodes. A continuing task does not naturally divide into identifiable episodes.
The Episode Lifecycle
An episode is one complete run from a starting point to an ending point. To trace an episodic task, follow three stages. The agent begins from a standard starting state or from a state selected from a standard distribution of starting states. The agent and environment then interact while the episode is in progress. Finally, the interaction reaches a special terminal state. After the terminal state is reached, the task resets and another episode can begin.
The State Sets S and S+
Episodic tasks use two related state sets. S contains all nonterminal states: states in which the episode has not yet ended. S+ contains all of the states in S together with the terminal state. The distinction matters because an agent can occupy states from S while an episode is in progress, but the terminal state must also be included when the complete state collection includes the episode boundary.
| State collection | What it contains | Role in an episode |
|---|---|---|
| S | All nonterminal states | States before the episode has ended |
| S+ | All states in S plus the terminal state | Includes the episode boundary |
Different Outcomes, One Boundary
A win and a loss
How can two different game outcomes use the same terminal state?
First play: The agent moves through nonterminal states in S, reaches the terminal state, and receives a reward associated with a win.
Second play: The agent again moves through nonterminal states in S, reaches the same terminal state, and receives a reward associated with a loss.
Episode boundary: The terminal state identifies where each play ends. The reward identifies the outcome, so the boundary and the outcome do not need to be represented by separate terminal states.
Reset: After either outcome, the task resets before another play begins.
Different outcomes can have different rewards while both episodes transition into one terminal state.
The terminal state represents where an episode ends. The reward can represent what happened at that ending, so different outcomes do not require different terminal states.
Interactions Without Episodes
A continuing task is an agent-environment interaction that does not naturally divide into identifiable episodes. The interaction continues without limit instead of reaching a natural point where one run ends and another begins.
| Feature | Episodic task | Continuing task |
|---|---|---|
| Organization | Naturally divides into separate episodes | Does not naturally divide into identifiable episodes |
| Ending | Reaches a terminal state | Has no natural point at which one interaction ends |
| After the ending | The task resets and another episode can begin | The interaction continues without limit |
Why Infinite Returns Cause Trouble
For a continuing task, the interaction continues without limit, so the final time step is represented as T = ∞. This creates a problem for the standard return formulation: if rewards continue to accumulate across an unlimited sequence of time steps, the return can be infinite.
A reward of +1 forever
What happens when a continuing task supplies a reward of +1 at every time step?
First steps: The interaction supplies +1 at the first time step, then +1 again at the next time step, and continues doing so.
No final step: Because the task is continuing, its final time step is represented as T = ∞.
Accumulation: There is no terminal point at which reward accumulation stops. The standard return formulation can therefore produce an infinite return.
Repeated +1 rewards across an interaction with T = ∞ illustrate why the standard return formulation becomes problematic for continuing tasks.
Common Classification Mistakes
Treating the terminal state as an ordinary middle-of-episode state.
The terminal state marks the end of the episode, and a reset follows.
Fix:
Separate the episode into its run toward the terminal state and the reset that begins the next episode.Confusing S with S+.
S contains only nonterminal states, while S+ also includes the terminal state.
Fix:
Use S for states before termination and S+ when the terminal state must be included.Assuming different outcomes require different terminal states.
Different rewards can represent different outcomes even when both episodes share one terminal state.
Fix:
Keep the episode boundary represented by the terminal state and use the reward to distinguish the outcome.Calling every long interaction continuing.
The defining issue is whether the interaction naturally divides into identifiable episodes, not merely how long one episode lasts.
Fix:
Look for a natural terminal point followed by a reset.Applying a finite-ending intuition to T = ∞.
A continuing task has no natural final time step, and the standard return formulation can produce an infinite return.
Fix:
Check whether the task continues without limit before reasoning about its return.
Classify the Interaction
A task begins from a starting state, proceeds through nonterminal states, reaches an outcome, and then resets before another run begins. Classify the task and identify the roles of S, S+, the terminal state, and the reset.
Hints
- Ask whether the interaction naturally divides into complete runs.
- Identify which states occur before the episode ends.
- Separate the episode boundary from the reward associated with its outcome.
A second task has no natural point at which one interaction ends and another begins. It continues without limit and supplies +1 at every time step. Classify the task and explain why the standard return formulation becomes problematic.
Hints
- Look for the absence of an identifiable terminal boundary.
- Represent the final time step as T = ∞.
- Consider what happens when +1 continues to accumulate without a final step.
Key Takeaways
- An episodic task naturally divides into separate episodes, each running from a starting state to a terminal state.
- Reaching the terminal state ends the episode; a reset then allows another episode to begin.
- S contains nonterminal states, while S+ contains those states together with the terminal state.
- Different outcomes can share one terminal state because rewards can distinguish the outcomes.
- A continuing task has no natural division into identifiable episodes, and T = ∞ can make the standard return formulation produce an infinite return.
Key Takeaways
- Episodic tasks have identifiable runs that end in a terminal state and then reset.
- S contains only nonterminal states; S+ also contains the terminal state.
- A shared terminal state can mark the end of different outcomes, while rewards distinguish those outcomes.
- Continuing tasks do not naturally divide into episodes and continue without limit.
- When T = ∞, repeated rewards such as +1 at every time step can make the standard return infinite.