Concepts / Episodic and Continuing Tasks

Episodic and Continuing Tasks

Episodic tasks have interactions that naturally divide into separate episodes.

  • Programming

One Run or an Ongoing Interaction

A reinforcement-learning interaction can have a clear ending, or it can continue without a natural point at which one interaction ends and another begins. This distinction produces two task types: episodic tasks and continuing tasks. The key question is whether the interaction naturally divides into complete runs with identifiable beginnings and endings.

An episodic task is organized into separate episodes. A continuing task does not naturally divide into identifiable episodes.

The Episode Lifecycle

An episode is one complete run from a starting point to an ending point. To trace an episodic task, follow three stages. The agent begins from a standard starting state or from a state selected from a standard distribution of starting states. The agent and environment then interact while the episode is in progress. Finally, the interaction reaches a special terminal state. After the terminal state is reached, the task resets and another episode can begin.

interactreach endpointresetnew episodeStarting stateepisode beginsNonterminal statesinteraction in progressTerminal stateepisode endsResetnext episode begins
What happens as an episodic interaction reaches its terminal state?

The State Sets S and S+

Episodic tasks use two related state sets. S contains all nonterminal states: states in which the episode has not yet ended. S+ contains all of the states in S together with the terminal state. The distinction matters because an agent can occupy states from S while an episode is in progress, but the terminal state must also be included when the complete state collection includes the episode boundary.

containscontainsalso containsSnonterminal statesNonterminal statesepisode continuesS+S and terminal stateTerminal stateepisode ends
Which states belong to S, and what additional state does S+ include?
State collectionWhat it containsRole in an episode
SAll nonterminal statesStates before the episode has ended
S+All states in S plus the terminal stateIncludes the episode boundary

Different Outcomes, One Boundary

A win and a loss

How can two different game outcomes use the same terminal state?

First play: The agent moves through nonterminal states in S, reaches the terminal state, and receives a reward associated with a win.

Second play: The agent again moves through nonterminal states in S, reaches the same terminal state, and receives a reward associated with a loss.

Episode boundary: The terminal state identifies where each play ends. The reward identifies the outcome, so the boundary and the outcome do not need to be represented by separate terminal states.

Reset: After either outcome, the task resets before another play begins.

Different outcomes can have different rewards while both episodes transition into one terminal state.

endsendsreward differsreward differsPlay 1nonterminal statesTerminal stateshared episode boundaryWin rewardoutcome of Play 1Play 2nonterminal statesLoss rewardoutcome of Play 2
How can a win and a loss use one terminal state while keeping their rewards different?

The terminal state represents where an episode ends. The reward can represent what happened at that ending, so different outcomes do not require different terminal states.

Interactions Without Episodes

A continuing task is an agent-environment interaction that does not naturally divide into identifiable episodes. The interaction continues without limit instead of reaching a natural point where one run ends and another begins.

reachesremains inEpisodic taskseparate episodesTerminal statereset followsContinuing taskno natural episodesOngoing interactioncontinues without limit
How does an interaction with identifiable episodes differ from one that continues indefinitely?
FeatureEpisodic taskContinuing task
OrganizationNaturally divides into separate episodesDoes not naturally divide into identifiable episodes
EndingReaches a terminal stateHas no natural point at which one interaction ends
After the endingThe task resets and another episode can beginThe interaction continues without limit

Why Infinite Returns Cause Trouble

For a continuing task, the interaction continues without limit, so the final time step is represented as T = ∞. This creates a problem for the standard return formulation: if rewards continue to accumulate across an unlimited sequence of time steps, the return can be infinite.

next stepnext stepcontinuesno endpointt = 0+1t = 1+1t = 2+1Further time steps+1 at each stepT = ∞no final step
How can receiving +1 at every time step make accumulated return grow without bound?

A reward of +1 forever

What happens when a continuing task supplies a reward of +1 at every time step?

First steps: The interaction supplies +1 at the first time step, then +1 again at the next time step, and continues doing so.

No final step: Because the task is continuing, its final time step is represented as T = ∞.

Accumulation: There is no terminal point at which reward accumulation stops. The standard return formulation can therefore produce an infinite return.

Repeated +1 rewards across an interaction with T = ∞ illustrate why the standard return formulation becomes problematic for continuing tasks.

Common Classification Mistakes

  • Treating the terminal state as an ordinary middle-of-episode state.

    The terminal state marks the end of the episode, and a reset follows.

    Fix: Separate the episode into its run toward the terminal state and the reset that begins the next episode.

  • Confusing S with S+.

    S contains only nonterminal states, while S+ also includes the terminal state.

    Fix: Use S for states before termination and S+ when the terminal state must be included.

  • Assuming different outcomes require different terminal states.

    Different rewards can represent different outcomes even when both episodes share one terminal state.

    Fix: Keep the episode boundary represented by the terminal state and use the reward to distinguish the outcome.

  • Calling every long interaction continuing.

    The defining issue is whether the interaction naturally divides into identifiable episodes, not merely how long one episode lasts.

    Fix: Look for a natural terminal point followed by a reset.

  • Applying a finite-ending intuition to T = ∞.

    A continuing task has no natural final time step, and the standard return formulation can produce an infinite return.

    Fix: Check whether the task continues without limit before reasoning about its return.

Classify the Interaction

EASY

A task begins from a starting state, proceeds through nonterminal states, reaches an outcome, and then resets before another run begins. Classify the task and identify the roles of S, S+, the terminal state, and the reset.

Hints
  • Ask whether the interaction naturally divides into complete runs.
  • Identify which states occur before the episode ends.
  • Separate the episode boundary from the reward associated with its outcome.
MEDIUM

A second task has no natural point at which one interaction ends and another begins. It continues without limit and supplies +1 at every time step. Classify the task and explain why the standard return formulation becomes problematic.

Hints
  • Look for the absence of an identifiable terminal boundary.
  • Represent the final time step as T = ∞.
  • Consider what happens when +1 continues to accumulate without a final step.

Key Takeaways

  1. An episodic task naturally divides into separate episodes, each running from a starting state to a terminal state.
  2. Reaching the terminal state ends the episode; a reset then allows another episode to begin.
  3. S contains nonterminal states, while S+ contains those states together with the terminal state.
  4. Different outcomes can share one terminal state because rewards can distinguish the outcomes.
  5. A continuing task has no natural division into identifiable episodes, and T = ∞ can make the standard return formulation produce an infinite return.

Key Takeaways

  • Episodic tasks have identifiable runs that end in a terminal state and then reset.
  • S contains only nonterminal states; S+ also contains the terminal state.
  • A shared terminal state can mark the end of different outcomes, while rewards distinguish those outcomes.
  • Continuing tasks do not naturally divide into episodes and continue without limit.
  • When T = ∞, repeated rewards such as +1 at every time step can make the standard return infinite.