Concepts / Continuing Tasks and On-Policy Distributions

Continuing Tasks and On-Policy Distributions

The on-policy distribution describes how state probability is distributed under a policy.

  • Programming

From Policy to State Visits

When reinforcement learning evaluates a policy, it does not necessarily treat every state as equally important. The useful question is: under this policy, which states are encountered, and how often? The on-policy distribution answers that question by assigning probabilities to states according to the behavior produced by the policy.

For an episodic task, the answer depends on the entire way an episode unfolds. Some visits happen immediately because the episode-start rule places the agent in a state. Other visits happen later because policy-controlled transitions move the agent into states reached from states already being visited. The on-policy distribution therefore reflects both the initial-state distribution and the transitions that occur while following the policy.

selectscausesproducescontributes toPolicyActionTransitionState visitsOn-policydistribution
How does a policy generate a distribution over states by selecting actions, causing transitions, and producing repeated state visits?

Tracing Expected Visits

A reliable way to calculate an episodic on-policy distribution is to follow the source of every visit. Begin with the episode-start rule. Record the expected visits supplied directly by starting episodes in each state. Then follow the policy-controlled transitions and add the visits that those transitions produce in later states.

Separating Starts from Later Transitions

Imagine an episodic setting with states A, B, and C. Suppose an episode-start rule supplies some visits directly to A and B. While following the policy, transitions from already visited states produce additional visits to B and C. How should the expected visits be reasoned about?

Begin with starts: Count the expected visits supplied by the initial-state distribution. These visits exist before any policy-following transition is considered.

Trace transitions: For each state already being visited, follow the policy-controlled transitions and add the expected visits that arrive in later states.

Combine sources: The expected visit count for a state includes visits received directly at episode start and visits received later through transitions.

Keep counts separate from probabilities: At this stage, the result is a collection of expected visit counts, not yet a probability distribution.

Expected visits arise from two sources: episode starts and transitions that occur while following the policy.

places episodes inplaces episodes inis followed byis followed byreachesreachescontributescontributescontributesEpisode startsinitial-state distributionState Aexpected visitsPolicy transitionslater visitsExpected visit countsall sources combinedState Bexpected visitsState Cexpected visits
How do episode starts and subsequent policy-driven transitions accumulate expected visits to each state over an episode?

Normalization and Initial States

Expected visit counts are not yet probabilities. A count tells you how much state occupancy was accumulated, while a probability describes a share of the total occupancy. To obtain the on-policy distribution, divide each state's expected visit count by the total expected number of visits across the states. The resulting probabilities sum to one.

StageWhat is being representedWhy it matters
Expected visitsHow often each state is expected to be visitedCaptures episode starts and policy-following transitions
Total expected visitsThe combined expected visits across statesProvides the quantity used for normalization
On-policy probabilitiesEach state's share of total expected visitsProduces probabilities that sum to one

The calculation has a visit-count stage followed by a probability stage.

countcountcountshareshareshareState Aexpected countNormalizationdivide by totalState AprobabilityState Bexpected countState BprobabilityState Cexpected countState Cprobability
Why do expected visit counts need to be divided by the total expected number of visits before they become state probabilities that sum to one?

Initial-state choices can change the final on-policy distribution even when the policy and transition behavior stay the same. If one state receives more episode starts, it can accumulate more expected visits immediately. A state that is rarely selected initially can still become common later if policy-following transitions lead into it.

changeschangesStart distribution1more starts in AExpected visitsdistribution 1Start distribution2more starts in BExpected visitsdistribution 2
How does changing the starting-state distribution alter the probability of visiting states later in an episode?

When Interaction Never Ends

A continuing task is an agent-environment interaction that does not naturally divide into identifiable episodes. The interaction continues without limit rather than reaching a natural point at which one interaction ends and another begins.

Interaction typeHow it is characterizedTime structure
Episodic taskBreaks naturally into identifiable episodesAn episode has a start and an interaction that ends
Continuing taskDoes not naturally divide into identifiable episodesAgent and environment continue interacting without limit

This distinction matters because a standard return formulation expects a final time step. In a continuing task, the final time step is represented as T = infinity. The interaction has no natural final step at which the usual finite ending of an episode can be applied.

reachescontinues asEpisodic taskidentifiable episodesEpisode endnew interaction beginsContinuing taskno natural episodesOngoing interactionno final step
What distinguishes an interaction that continues indefinitely from one that naturally reaches an identifiable terminal state and resets?

The Infinite-Return Problem

When T is infinity, the standard return formulation can produce an infinite return. The clearest illustration is a continuing task in which the agent receives +1 at every time step. Because the interaction continues without limit and the positive reward keeps arriving, the accumulated return is infinite.

next stepnext stepcontinuesno final stepT = 0+1T = 1+1T = 2+1More time steps+1 repeatedlyT = infinityinfinite return
What happens to the return when there is no final time step and a reward such as +1 is received at every time step?

Reasoning Checks

  • Treating every state as equally important under a policy.

    The distribution describes which states are encountered under the policy and how often. State frequencies depend on episode starts and policy-following transitions.

    Fix: Trace the visits produced by the initial-state distribution and the transitions generated while following the policy.

  • Ignoring the initial-state distribution.

    For episodic tasks, initial-state selection is part of the distribution calculation. Changing the starting-state distribution can change expected visits.

    Fix: Begin the calculation with the expected visits supplied directly by episode starts.

  • Calling expected visit counts probabilities before normalization.

    Counts represent accumulated visits, not shares of total visits. They do not automatically sum to one.

    Fix: Add the expected visits across states, then divide each state's count by that total.

  • Counting only initial visits.

    Expected visits arise from both episode starts and policy-following transitions.

    Fix: Trace the transitions after each start and include the later visits they produce.

  • Assuming every interaction has a finite final time step.

    Continuing tasks do not naturally divide into identifiable episodes, and their final time step is represented as T = infinity.

    Fix: First determine whether the interaction has identifiable episodes. If it is continuing, examine whether the return can become infinite.

MEDIUM

A policy is evaluated in an episodic task. The initial-state rule is changed, but the policy and transition behavior remain the same. Explain why the on-policy distribution can still change. Then describe the two stages you would use to calculate the new distribution.

Hints
  • Identify what changes before the first transition occurs.
  • Separate expected visit counts from normalized probabilities.
  • Remember that later transitions can carry the effect of a changed starting distribution into other states.
EASY

Classify the following situation conceptually: an agent and environment interact without a natural point at which one interaction ends and another begins, while the agent receives +1 at every time step. Identify the task type and explain the return problem.

Hints
  • Use the distinction between identifiable episodes and interactions that continue without limit.
  • Consider what happens when the final time step is T = infinity.
  • Ask whether the repeated positive rewards can accumulate to a finite total.

A Reliable Solution Sequence

  1. Identify whether the interaction is episodic or continuing.
  2. For an episodic on-policy distribution, write down how episode starts place visits into states.
  3. Trace the policy-following transitions and add the later expected visits they produce.
  4. Keep the resulting expected visit counts separate from probabilities.
  5. Normalize the counts by the total expected number of visits so the state probabilities sum to one.
  6. For a continuing task, check whether the absence of a final time step makes the standard return infinite, especially when positive rewards repeat indefinitely.

The on-policy distribution describes state occupancy under a policy: which states the policy causes the agent to encounter and how often.

In episodic tasks, initial-state selection and policy-following transitions jointly determine expected visits. Normalization converts those expected counts into probabilities.

Continuing tasks do not naturally divide into identifiable episodes. Because their final time step is represented as T = infinity, repeated rewards such as +1 at every time step can produce an infinite return.

Key Takeaways

  • The on-policy distribution assigns state probabilities according to the states encountered under a policy and how often they are encountered.
  • For episodic tasks, expected visits come from both the initial-state distribution and policy-following transitions.
  • Expected visit counts must be normalized by the total expected visits before they represent probabilities that sum to one.
  • Changing only the initial-state distribution can change the on-policy distribution.
  • Continuing tasks have no natural episode boundary, so a standard return can become infinite when positive rewards repeat indefinitely.