Policy Gradient for Continuing Problems
Continuing problems have no episode endpoints, so policy performance is measured by average reward per time step over the long run.
One Uninterrupted Stream
Many reinforcement-learning problems are organized into episodes. An agent starts, receives rewards, and eventually reaches an endpoint. A continuing problem has no such episode boundary: the agent keeps acting in one uninterrupted stream. That difference changes how policy performance is described. Instead of using the total reward from one episode, we measure the reward obtained per time step over the long run.
The central question is not how much reward one completed episode produced. It is how much reward per time step the policy produces as the continuing process runs for a very long time.
The Average Reward Rate
The average reward rate is the limiting average reward as time approaches infinity. To form the underlying average, consider the rewards collected over an increasingly long prefix of the continuing stream and divide by the number of time steps in that prefix. The long-run reward rate is the value that this average approaches as the stream continues indefinitely.
A Continuing Reward Stream
A policy produces the reward sequence 2, 0, 1, 2, 0, 1, and continues in the same pattern. What does the average reward over a finite prefix measure, and what does the long-run average reward rate describe?
Observe a finite prefix: The first three rewards have an average of 1. This is a calculation based only on rewards that have already been observed.
Extend the stream: If the process continues, the average can be recomputed over six rewards, nine rewards, and increasingly longer prefixes.
Separate the meanings: Each finite calculation is a sample average for a particular prefix. The continuing-problem performance measure is the long-run average reward rate approached as the number of time steps grows without bound.
A finite sample average describes an observed prefix. The average reward rate is the limiting long-run measure used to evaluate the policy in a continuing problem.
State Frequencies Under a Policy
To understand the long-run reward rate, first track where the agent spends its time. Under policy π, the probability of being in a state can settle into a long-run distribution. This distribution is called the steady-state distribution and is written as dπ.
The distribution dπ summarizes the long-run state frequencies produced by policy π. A policy therefore affects long-run performance in two connected ways: it determines how the agent acts, and those actions determine how often the continuing process occupies different states. Understanding those state frequencies helps explain how rewards accumulate over time.
Generated example: Imagine a continuing system with three states. Under one policy, the process spends most of its time in a high-reward state; under another policy, it spends more time in lower-reward states. Even without episode endpoints, comparing the policies can focus on the long-run reward per time step produced by their different state frequencies.
Why Ergodicity Matters
The steady-state idea requires an additional assumption. The limiting distribution must not depend on the initial state S0. This is the ergodicity assumption in this context: different starting states are assumed to lead to the same long-run distribution when the agent follows policy π.
Independence from the initial state matters because it gives the policy a single long-run state distribution for this analysis. If the limiting distribution depended on where the agent began, then the same policy could have different long-run state frequencies for different starting states. The steady-state description would no longer represent one common long-run outcome of the policy.
Finite Prefixes and Limiting Performance
Suppose the beginning of a continuing reward stream contains unusually high or unusually low rewards. The average over that finite prefix can differ substantially from the value approached over a much longer stream. That difference is not a contradiction. The finite average summarizes what has already been observed, while the long-run average reward rate is defined by the limiting behavior as time approaches infinity.
What do you think happens?
A continuing process begins with several high rewards and later settles into a different pattern. Which statement is correct?
Reveal answer
Answer: The initial finite average may differ from the long-run average reward rate.
A finite average uses only a chosen observed prefix. The average reward rate is the limiting average reward as time approaches infinity.
Mistakes in Long-Run Reasoning
Treating a continuing problem as if it must be divided into episodes
A continuing problem has no episode endpoint. Its performance is described per time step over the long run.
Fix:
Use the average reward rate, defined by the limiting average reward as time approaches infinity.Calling one finite sample average the long-run reward rate
A finite calculation summarizes an observed prefix, whereas the long-run criterion concerns the value approached as the process continues indefinitely.
Fix:
Keep the finite sample average and the limiting average reward rate conceptually separate.Ignoring the state distribution induced by the policy
The steady-state distribution dπ describes the limiting distribution of states under policy π and helps explain long-run behavior.
Fix:
Track the long-run state frequencies under the policy.Assuming the initial state can be ignored without an assumption
The common steady-state interpretation relies on the ergodicity assumption that the limiting distribution is independent of S0.
Fix:
State the ergodicity assumption when using a steady-state distribution as a policy-level long-run description.
Check Your Understanding
Explain in your own words why the average of rewards from a finite prefix is not automatically the average reward rate for a continuing problem. Then explain what dπ represents and state the role of the ergodicity assumption.
Hints
- Start by identifying whether the quantity uses a fixed, already observed number of time steps or a limit as time approaches infinity.
- Describe dπ as a limiting distribution of states under policy π.
- Mention that ergodicity makes this limiting distribution independent of the initial state S0.
Classifying Three Statements
Classify each statement as describing a finite sample average, a long-run average reward rate, or the steady-state distribution.
Statement A: The average of rewards collected during the first 20 time steps is a finite sample average.
Statement B: The reward per time step approached as the continuing process runs indefinitely is the average reward rate.
Statement C: The limiting distribution of states under policy π is the steady-state distribution dπ.
Finite sample average describes an observed prefix; average reward rate describes limiting reward per time step; dπ describes limiting state frequencies under the policy.
Long-Run Policy Evaluation
- A continuing problem has no episode endpoints, so policy performance is measured per time step over the long run.
- The average reward rate is the limiting average reward as time approaches infinity.
- The steady-state distribution dπ is the limiting distribution of states under policy π.
- The ergodicity assumption makes the limiting state distribution independent of the initial state S0.
- A finite sample average summarizes an observed prefix and must not be confused with the long-run average reward rate.
Key Takeaways
- Continuing problems evaluate a policy through reward per time step rather than total reward from a completed episode.
- The average reward rate is the limiting average reward as time approaches infinity.
- The steady-state distribution dπ records the limiting distribution of states under a policy.
- Ergodicity ensures that this limiting distribution does not depend on the initial state.
- A finite prefix average is an observation-based calculation, not automatically the long-run performance measure.