Concepts / Steady-State Distribution

Steady-State Distribution

Continuing problems have no episode endpoints, so policy performance is measured by average reward per time step over the long run.

  • Programming

Why Continuing Tasks Need a Different Measure

Many reinforcement-learning problems are organized into episodes: the agent starts, receives rewards, and eventually reaches an endpoint. A continuing problem has no such episode boundary. The agent keeps acting, so evaluating a policy cannot depend on the total reward from one completed episode. Instead, performance is described by the reward obtained per time step over the long run.

The average reward rate is the limiting average reward per time step as time approaches infinity.

Following State Occupancy Over Time

To understand the long-run reward rate, first track where the agent spends its time. Under policy π, the probability of being in a state can change during the early steps. The process may begin from a particular state, and the initial distribution can strongly influence what is observed at first. The important question is what remains after the process has continued for a long time.

policy and transitions acttime growsInitial distributionDepends on where the MDPstartsChanging distributionEarly state probabilitiesSteady-statedistribution dπLong-run stateprobabilities
What changes during the early steps of a continuing task, and what remains after temporary starting effects disappear?

The changing distribution seen during the first few steps is not itself the steady-state distribution. The steady-state distribution is the limiting distribution reached as time grows, when that limit exists. It describes long-run state occupancy rather than the temporary pattern immediately after the task begins.

Defining the Steady-State Distribution

The steady-state distribution dπ is the limiting distribution of states when the agent follows policy π.

For a particular state s, the definition asks for the probability that the process is in s at time t after the earlier actions were selected according to policy π, and then considers what that probability approaches as t grows without bound. The notation combines a state, a policy that governs action selection, and a long-run limit.

governsinfluenceproducesettle intoPolicy πSelects actionsActionsSelected over timeTransitionsMove the process betweenstatesState visitsLong-run occupancydπSteady-state distribution
How are long-run state-visit proportions determined by the policy and the transitions between states?

The long-run state expectation is determined by both the policy and the transition probabilities. Once the process is in the steady-state distribution, continuing to select actions according to π preserves that distribution. In other words, applying the policy does not change the distribution as a distribution, even though individual time steps can still involve state transitions.

Tracing Different Starting States

Two Starts, One Long-Run Distribution

Imagine two runs of the same continuing MDP. In the first run, the process begins in one state. In the second run, it begins in a different state. Both runs then select actions according to the same policy π.

Step 1: Compare the beginning: The state distributions can differ because the two runs began from different initial states. Early observations therefore retain information about how each run started.

Step 2: Continue applying π: The policy and the MDP transitions continue to determine how states are visited. The distributions may change as the processes move away from their initial conditions.

Step 3: Consider the long run: Under the ergodicity assumption, the limiting distribution is independent of the initial state. Both runs therefore approach the same steady-state distribution dπ.

Step 4: Interpret early decisions: Early decisions can affect the temporary sequence of states, but their influence is temporary in the long-run description provided by the steady-state distribution.

The two runs can look different at first while having the same long-run state distribution under policy π.

policy actslong runpolicy actslong runInitial state ARun 1Initial state BRun 2Early distribution ATemporary effectEarly distribution BTemporary effectdπCommon limitingdistribution
How do state distributions starting from different initial states gradually converge to the same long-run distribution under a policy?

Ergodicity and Temporary Effects

In this context, ergodicity is the assumption that the limiting state distribution under a policy exists and is independent of the initial state.

Ergodicity separates temporary starting effects from long-run behavior. At early times, the state distribution may depend strongly on the initial state and on actions selected early by the agent. In an ergodic MDP, those influences do not determine the limiting distribution. As the process continues, the long-run expectation of being in a state is determined by the policy and the MDP transition probabilities.

influences early stepstime continueslimit existsInitial stateStarting conditionTemporary behaviorEarly state probabilitiesErgodic long-runbehaviorLimit independent of startdπStable limitingdistribution
What connectivity and recurrence behavior allow a policy to produce a stable long-run distribution that does not depend on where the process began?

Finite Averages and Long-Run Rates

QuantityWhat it usesWhat it means
Finite sample averageRewards observed during a limited runThe average reward measured so far
Long-run average reward rateThe limiting average as time approaches infinityThe continuing-problem criterion for policy performance

Why Early Rewards Can Mislead

Suppose a continuing task is observed for only a short initial portion of its operation. The rewards during that portion are averaged, and the result is compared with the policy's long-run average reward rate.

Finite calculation: The observed average summarizes only the rewards that have already appeared in the limited sample.

Long-run interpretation: The average reward rate is defined by the limiting average as the run becomes indefinitely long, not by the temporary average from the initial portion.

Reason for the difference: Early state distributions can depend on the initial state and early decisions. Those temporary conditions can affect the finite sample without determining the long-run rate in an ergodic MDP.

A finite sample average can differ from the long-run average reward rate because the sample may still contain temporary starting effects.

can includeis defined byFinite sampleaverageLimited observed rewardsAverage reward rateLimiting averageTemporary effectsMay affect the finitesampleTime approachesinfinityDefines the long-runcriterion
How can the average reward measured over a short observed run differ from the reward rate approached as the run becomes indefinitely long?

Common Interpretation Mistakes

  • Treating the first few state distributions as the steady-state distribution.

    Early distributions can depend strongly on the initial state and early decisions.

    Fix: Identify dπ as the limiting distribution reached as time grows under policy π.

  • Calling a finite sample average the long-run average reward rate.

    A finite sample averages rewards that have already been observed, while the long-run rate is defined by the limiting average as time approaches infinity.

    Fix: Keep the finite calculation separate from the long-run performance definition.

  • Assuming independence from the initial state means the initial state has no early effect.

    Ergodicity concerns the limiting distribution, not equality during every early step.

    Fix: Describe the initial-state influence as temporary and the limiting distribution as independent of the initial state.

  • Defining the steady-state distribution without mentioning the policy.

    The long-run state expectation is determined by the policy together with the transition probabilities.

    Fix: Always interpret dπ as the limiting state distribution under policy π.

Check Your Understanding

EASY

A continuing MDP is run under policy π from two different initial states. During the first few steps, the state distributions differ. Later, both runs approach the same limiting distribution. What concept explains this result, and what is the name of the common limiting distribution?

Hints
  • Focus on the assumption that the limiting distribution does not depend on the initial state.
  • The distribution is named using the policy π.

What do you think happens?

If two runs begin from different states but follow the same policy in an ergodic MDP, should their early state distributions necessarily be identical?

  • Yes, because the same policy is used
  • No, they can differ temporarily before approaching the same limiting distribution
  • Yes, because the steady-state distribution applies immediately
  • No, because an ergodic MDP has no limiting distribution
Reveal answer

Answer: No, they can differ temporarily before approaching the same limiting distribution.

Ergodicity makes the limiting distribution independent of the initial state. It does not require the distributions to match during the early steps.

Key Takeaways

  1. A continuing problem has no episode endpoints, so policy performance is measured using average reward per time step over the long run.
  2. The average reward rate is the limiting average reward as time approaches infinity.
  3. The steady-state distribution dπ is the limiting distribution of states under policy π.
  4. Ergodicity means that this limiting distribution exists and is independent of the initial state.
  5. Initial conditions and early decisions can affect temporary behavior without determining the long-run distribution in an ergodic MDP.

Key Takeaways

  • Continuing tasks require a long-run average reward rate because they do not have episode boundaries.
  • The steady-state distribution dπ describes the limiting state probabilities under policy π.
  • The policy and transition probabilities determine the long-run state expectation.
  • Ergodicity makes the limiting distribution independent of the initial state.
  • Finite observations can reflect temporary early effects and should not be confused with the long-run performance criterion.