Steady-State Distribution
Continuing problems have no episode endpoints, so policy performance is measured by average reward per time step over the long run.
Why Continuing Tasks Need a Different Measure
Many reinforcement-learning problems are organized into episodes: the agent starts, receives rewards, and eventually reaches an endpoint. A continuing problem has no such episode boundary. The agent keeps acting, so evaluating a policy cannot depend on the total reward from one completed episode. Instead, performance is described by the reward obtained per time step over the long run.
The average reward rate is the limiting average reward per time step as time approaches infinity.
Following State Occupancy Over Time
To understand the long-run reward rate, first track where the agent spends its time. Under policy π, the probability of being in a state can change during the early steps. The process may begin from a particular state, and the initial distribution can strongly influence what is observed at first. The important question is what remains after the process has continued for a long time.
The changing distribution seen during the first few steps is not itself the steady-state distribution. The steady-state distribution is the limiting distribution reached as time grows, when that limit exists. It describes long-run state occupancy rather than the temporary pattern immediately after the task begins.
Defining the Steady-State Distribution
The steady-state distribution dπ is the limiting distribution of states when the agent follows policy π.
For a particular state s, the definition asks for the probability that the process is in s at time t after the earlier actions were selected according to policy π, and then considers what that probability approaches as t grows without bound. The notation combines a state, a policy that governs action selection, and a long-run limit.
The long-run state expectation is determined by both the policy and the transition probabilities. Once the process is in the steady-state distribution, continuing to select actions according to π preserves that distribution. In other words, applying the policy does not change the distribution as a distribution, even though individual time steps can still involve state transitions.
Tracing Different Starting States
Two Starts, One Long-Run Distribution
Imagine two runs of the same continuing MDP. In the first run, the process begins in one state. In the second run, it begins in a different state. Both runs then select actions according to the same policy π.
Step 1: Compare the beginning: The state distributions can differ because the two runs began from different initial states. Early observations therefore retain information about how each run started.
Step 2: Continue applying π: The policy and the MDP transitions continue to determine how states are visited. The distributions may change as the processes move away from their initial conditions.
Step 3: Consider the long run: Under the ergodicity assumption, the limiting distribution is independent of the initial state. Both runs therefore approach the same steady-state distribution dπ.
Step 4: Interpret early decisions: Early decisions can affect the temporary sequence of states, but their influence is temporary in the long-run description provided by the steady-state distribution.
The two runs can look different at first while having the same long-run state distribution under policy π.
Ergodicity and Temporary Effects
In this context, ergodicity is the assumption that the limiting state distribution under a policy exists and is independent of the initial state.
Ergodicity separates temporary starting effects from long-run behavior. At early times, the state distribution may depend strongly on the initial state and on actions selected early by the agent. In an ergodic MDP, those influences do not determine the limiting distribution. As the process continues, the long-run expectation of being in a state is determined by the policy and the MDP transition probabilities.
Finite Averages and Long-Run Rates
| Quantity | What it uses | What it means |
|---|---|---|
| Finite sample average | Rewards observed during a limited run | The average reward measured so far |
| Long-run average reward rate | The limiting average as time approaches infinity | The continuing-problem criterion for policy performance |
Why Early Rewards Can Mislead
Suppose a continuing task is observed for only a short initial portion of its operation. The rewards during that portion are averaged, and the result is compared with the policy's long-run average reward rate.
Finite calculation: The observed average summarizes only the rewards that have already appeared in the limited sample.
Long-run interpretation: The average reward rate is defined by the limiting average as the run becomes indefinitely long, not by the temporary average from the initial portion.
Reason for the difference: Early state distributions can depend on the initial state and early decisions. Those temporary conditions can affect the finite sample without determining the long-run rate in an ergodic MDP.
A finite sample average can differ from the long-run average reward rate because the sample may still contain temporary starting effects.
Common Interpretation Mistakes
Treating the first few state distributions as the steady-state distribution.
Early distributions can depend strongly on the initial state and early decisions.
Fix:
Identify dπ as the limiting distribution reached as time grows under policy π.Calling a finite sample average the long-run average reward rate.
A finite sample averages rewards that have already been observed, while the long-run rate is defined by the limiting average as time approaches infinity.
Fix:
Keep the finite calculation separate from the long-run performance definition.Assuming independence from the initial state means the initial state has no early effect.
Ergodicity concerns the limiting distribution, not equality during every early step.
Fix:
Describe the initial-state influence as temporary and the limiting distribution as independent of the initial state.Defining the steady-state distribution without mentioning the policy.
The long-run state expectation is determined by the policy together with the transition probabilities.
Fix:
Always interpret dπ as the limiting state distribution under policy π.
Check Your Understanding
A continuing MDP is run under policy π from two different initial states. During the first few steps, the state distributions differ. Later, both runs approach the same limiting distribution. What concept explains this result, and what is the name of the common limiting distribution?
Hints
- Focus on the assumption that the limiting distribution does not depend on the initial state.
- The distribution is named using the policy π.
What do you think happens?
If two runs begin from different states but follow the same policy in an ergodic MDP, should their early state distributions necessarily be identical?
Reveal answer
Answer: No, they can differ temporarily before approaching the same limiting distribution.
Ergodicity makes the limiting distribution independent of the initial state. It does not require the distributions to match during the early steps.
Key Takeaways
- A continuing problem has no episode endpoints, so policy performance is measured using average reward per time step over the long run.
- The average reward rate is the limiting average reward as time approaches infinity.
- The steady-state distribution dπ is the limiting distribution of states under policy π.
- Ergodicity means that this limiting distribution exists and is independent of the initial state.
- Initial conditions and early decisions can affect temporary behavior without determining the long-run distribution in an ergodic MDP.
Key Takeaways
- Continuing tasks require a long-run average reward rate because they do not have episode boundaries.
- The steady-state distribution dπ describes the limiting state probabilities under policy π.
- The policy and transition probabilities determine the long-run state expectation.
- Ergodicity makes the limiting distribution independent of the initial state.
- Finite observations can reflect temporary early effects and should not be confused with the long-run performance criterion.