Concepts / Policy Gradient for Continuing Problems

Policy Gradient for Continuing Problems

Continuing problems have no episode endpoints, so policy performance is measured by average reward per time step over the long run.

  • Programming

One Uninterrupted Stream

Many reinforcement-learning problems are organized into episodes. An agent starts, receives rewards, and eventually reaches an endpoint. A continuing problem has no such episode boundary: the agent keeps acting in one uninterrupted stream. That difference changes how policy performance is described. Instead of using the total reward from one episode, we measure the reward obtained per time step over the long run.

actscontinuescontinuesno endpointStartstate and reward streamTime step 1reward observedTime step 2reward observedTime step 3reward observedLater timestream continues
How does reward measurement change when experience is one uninterrupted stream rather than a sequence of episodes with terminal states?

The central question is not how much reward one completed episode produced. It is how much reward per time step the policy produces as the continuing process runs for a very long time.

The Average Reward Rate

The average reward rate is the limiting average reward as time approaches infinity. To form the underlying average, consider the rewards collected over an increasingly long prefix of the continuing stream and divide by the number of time steps in that prefix. The long-run reward rate is the value that this average approaches as the stream continues indefinitely.

A Continuing Reward Stream

A policy produces the reward sequence 2, 0, 1, 2, 0, 1, and continues in the same pattern. What does the average reward over a finite prefix measure, and what does the long-run average reward rate describe?

Observe a finite prefix: The first three rewards have an average of 1. This is a calculation based only on rewards that have already been observed.

Extend the stream: If the process continues, the average can be recomputed over six rewards, nine rewards, and increasingly longer prefixes.

Separate the meanings: Each finite calculation is a sample average for a particular prefix. The continuing-problem performance measure is the long-run average reward rate approached as the number of time steps grows without bound.

A finite sample average describes an observed prefix. The average reward rate is the limiting long-run measure used to evaluate the policy in a continuing problem.

State Frequencies Under a Policy

To understand the long-run reward rate, first track where the agent spends its time. Under policy π, the probability of being in a state can settle into a long-run distribution. This distribution is called the steady-state distribution and is written as dπ.

guidesvisitsvisitsvisitscontributescontributescontributesPolicy πchooses actionsContinuing processkeeps actingState Afrequency in dπState Bfrequency in dπState Cfrequency in dπAverage reward ratelong-run reward per timestep
How does a policy determine the long-run proportion of time spent in each state, and how do those state frequencies contribute to average reward?

The distribution dπ summarizes the long-run state frequencies produced by policy π. A policy therefore affects long-run performance in two connected ways: it determines how the agent acts, and those actions determine how often the continuing process occupies different states. Understanding those state frequencies helps explain how rewards accumulate over time.

Generated example: Imagine a continuing system with three states. Under one policy, the process spends most of its time in a high-reward state; under another policy, it spends more time in lower-reward states. Even without episode endpoints, comparing the policies can focus on the long-run reward per time step produced by their different state frequencies.

Why Ergodicity Matters

The steady-state idea requires an additional assumption. The limiting distribution must not depend on the initial state S0. This is the ergodicity assumption in this context: different starting states are assumed to lead to the same long-run distribution when the agent follows policy π.

follows πfollows πfollows πconverges toconverges toconverges toInitial state AS0 = AInitial state BS0 = BInitial state CS0 = CEvolving processfollows πEvolving processfollows πEvolving processfollows πSteady-statedistribution dπindependent of S0
How can a continuing system eventually produce the same long-run reward rate regardless of where it started?

Independence from the initial state matters because it gives the policy a single long-run state distribution for this analysis. If the limiting distribution depended on where the agent began, then the same policy could have different long-run state frequencies for different starting states. The steady-state description would no longer represent one common long-run outcome of the policy.

Finite Prefixes and Limiting Performance

longer prefixes reveal the limitFinite sampleaveragecomputed from observedrewardsLong-run averagereward ratelimiting average as timegrows
How can the average reward computed over a finite prefix differ from the value approached as the continuing process runs indefinitely?

Suppose the beginning of a continuing reward stream contains unusually high or unusually low rewards. The average over that finite prefix can differ substantially from the value approached over a much longer stream. That difference is not a contradiction. The finite average summarizes what has already been observed, while the long-run average reward rate is defined by the limiting behavior as time approaches infinity.

What do you think happens?

A continuing process begins with several high rewards and later settles into a different pattern. Which statement is correct?

  • The initial finite average is automatically the long-run average reward rate.
  • The initial finite average may differ from the long-run average reward rate.
  • A continuing problem cannot be evaluated because it has no endpoint.
  • The initial state must determine the limiting distribution.
Reveal answer

Answer: The initial finite average may differ from the long-run average reward rate.

A finite average uses only a chosen observed prefix. The average reward rate is the limiting average reward as time approaches infinity.

Mistakes in Long-Run Reasoning

  • Treating a continuing problem as if it must be divided into episodes

    A continuing problem has no episode endpoint. Its performance is described per time step over the long run.

    Fix: Use the average reward rate, defined by the limiting average reward as time approaches infinity.

  • Calling one finite sample average the long-run reward rate

    A finite calculation summarizes an observed prefix, whereas the long-run criterion concerns the value approached as the process continues indefinitely.

    Fix: Keep the finite sample average and the limiting average reward rate conceptually separate.

  • Ignoring the state distribution induced by the policy

    The steady-state distribution dπ describes the limiting distribution of states under policy π and helps explain long-run behavior.

    Fix: Track the long-run state frequencies under the policy.

  • Assuming the initial state can be ignored without an assumption

    The common steady-state interpretation relies on the ergodicity assumption that the limiting distribution is independent of S0.

    Fix: State the ergodicity assumption when using a steady-state distribution as a policy-level long-run description.

Check Your Understanding

MEDIUM

Explain in your own words why the average of rewards from a finite prefix is not automatically the average reward rate for a continuing problem. Then explain what dπ represents and state the role of the ergodicity assumption.

Hints
  • Start by identifying whether the quantity uses a fixed, already observed number of time steps or a limit as time approaches infinity.
  • Describe dπ as a limiting distribution of states under policy π.
  • Mention that ergodicity makes this limiting distribution independent of the initial state S0.

Classifying Three Statements

Classify each statement as describing a finite sample average, a long-run average reward rate, or the steady-state distribution.

Statement A: The average of rewards collected during the first 20 time steps is a finite sample average.

Statement B: The reward per time step approached as the continuing process runs indefinitely is the average reward rate.

Statement C: The limiting distribution of states under policy π is the steady-state distribution dπ.

Finite sample average describes an observed prefix; average reward rate describes limiting reward per time step; dπ describes limiting state frequencies under the policy.

Long-Run Policy Evaluation

  1. A continuing problem has no episode endpoints, so policy performance is measured per time step over the long run.
  2. The average reward rate is the limiting average reward as time approaches infinity.
  3. The steady-state distribution dπ is the limiting distribution of states under policy π.
  4. The ergodicity assumption makes the limiting state distribution independent of the initial state S0.
  5. A finite sample average summarizes an observed prefix and must not be confused with the long-run average reward rate.

Key Takeaways

  • Continuing problems evaluate a policy through reward per time step rather than total reward from a completed episode.
  • The average reward rate is the limiting average reward as time approaches infinity.
  • The steady-state distribution dπ records the limiting distribution of states under a policy.
  • Ergodicity ensures that this limiting distribution does not depend on the initial state.
  • A finite prefix average is an observation-based calculation, not automatically the long-run performance measure.