Episodic Tasks
Continuing tasks describe agent-environment interactions that do not naturally divide into identifiable episodes.
When Interaction Has No Natural End
An interaction between an agent and an environment can have either of two broad structures. It may continue without a natural point at which one interaction ends and another begins. Or it may divide naturally into separate, identifiable episodes. This distinction affects how we describe time, termination, rewards, and the return.
Continuing Tasks and Infinite Time
A continuing task is an agent-environment interaction that does not naturally divide into identifiable episodes. The agent and environment continue interacting without limit rather than reaching a natural boundary between one interaction and the next.
In a continuing task, the final time step is represented as T = ∞ because the interaction has no limiting final step. This creates a problem for the standard return formulation: under these conditions, the return can be infinite rather than a finite quantity.
Why Repeated Rewards Can Diverge
A reward of +1 at every time step
Consider a continuing task in which the agent receives +1 at every time step.
First step: The interaction produces a reward of +1.
More steps: The interaction continues, and each additional time step produces another reward of +1.
No final step: Because the continuing task has T = ∞, there is no finite endpoint at which the sequence of rewards stops.
Return consequence: The standard return formulation can therefore produce an infinite return.
Repeated rewards of +1 over an interaction that continues without limit illustrate how the return can become infinite.
The issue is not that a single reward of +1 is unusually large. The issue is that the interaction supplies another +1 at every time step and never reaches a finite final time step.
Episode Boundaries and Reset
An episodic task is a reinforcement-learning interaction that naturally divides into separate episodes. Each episode is one complete run from a starting point to an ending point.
To trace an episodic task, follow three parts. The agent begins from a standard starting state or from a state selected from a standard distribution of starting states. The agent and environment then interact while the episode is in progress. Finally, the interaction reaches a special terminal state. After the terminal state, the task resets and another episode can begin.
State Sets Before and At Termination
Episodic tasks use two related state sets. S contains all nonterminal states: states in which the episode has not yet ended. S+ contains all of the states in S together with the terminal state. Thus, S names the states before termination, while S+ is the larger state collection used when the terminal state must also be included.
| State collection | What it contains | When it is useful |
|---|---|---|
| S | All nonterminal states | When describing states before the episode has ended |
| S+ | All nonterminal states together with the terminal state | When the terminal state must be included in the state collection |
The two related state sets used for episodic tasks.
One Terminal State, Different Outcomes
Imagine two plays of a game. The first play ends with a win and the second ends with a loss. These outcomes differ, so their rewards may differ. Nevertheless, both plays can be described as moving through nonterminal states in S, reaching the terminal state, and then resetting before the next play.
The terminal state represents the episode boundary, not every detail of the outcome. Different rewards can represent different outcomes even when both outcomes lead to the same terminal state. This separates two ideas: the terminal state says that the episode has ended, while the reward can distinguish what happened at that ending.
Mistakes About Episodes
Treating every interaction as one unbroken stream
Some interactions have a natural stopping point and can be divided into separate episodes.
Fix:
Identify the terminal state and treat the reset as the beginning of the next episode.Assuming that a terminal state is just another nonterminal state
S contains nonterminal states, whereas the terminal state is included only when using the larger set S+.
Fix:
Use S for nonterminal states and S+ when the terminal state must also be included.Assuming different outcomes require different terminal states
Different outcomes can have different rewards while sharing one terminal state.
Fix:
Use the terminal state to mark the episode boundary and the reward to distinguish the outcome.Ignoring the consequence of T = ∞ in a continuing task
A continuing task has no finite final time step, so the standard return formulation can produce an infinite return.
Fix:
Check whether the interaction has a natural endpoint before applying an episodic interpretation.
Classifying an Interaction
A task begins from a starting state, proceeds through nonterminal states, reaches a special terminal state, resets, and then begins another run. Classify the task as continuing or episodic. Then identify which state set contains the terminal state: S or S+.
Hints
- Look for a natural endpoint and a reset.
- S contains only nonterminal states.
- S+ includes the terminal state as well.
What do you think happens?
A game has one terminal state. One play ends in a win and another ends in a loss. Must the game use two terminal states?
Reveal answer
Answer: No, because different rewards can distinguish the outcomes.
The terminal state marks that the episode has ended. Different rewards may represent the different outcomes even when both plays share that terminal state.
Key Takeaways
- A continuing task has no natural division into identifiable episodes and continues without limit, so its final time step is represented as T = ∞.
- The standard return formulation can produce an infinite return in a continuing task, especially when the agent receives +1 at every time step.
- An episodic task has a natural endpoint: the agent moves from a starting state through nonterminal states to a terminal state, followed by a reset.
- S contains nonterminal states, while S+ contains those states plus the terminal state.
- Different outcomes can share one terminal state because rewards can distinguish the outcomes while the terminal state marks the episode boundary.
Key Takeaways
- Continuing tasks continue without limit and do not naturally divide into episodes.
- When T = ∞, repeated rewards can make the standard return infinite.
- Episodic tasks move from a starting state through nonterminal states to a terminal state, then reset.
- S contains nonterminal states; S+ also includes the terminal state.
- A shared terminal state can mark the end of different outcomes whose rewards differ.