Concepts / Long-Run State Distributions

Long-Run State Distributions

Ergodicity separates temporary starting effects from long-run behavior.

  • Programming

The Long-Run Question

When a continuing task begins, the states observed during its first few steps can depend strongly on the initial state and on the actions selected early by the agent. Long-run analysis asks a different question: after the process has continued for a long time, what determines the probability of being in each state? In an ergodic Markov decision process, the influence of the starting state and early decisions is temporary. The long-run state expectation is determined by the policy and the MDP transition probabilities.

Ergodicity separates temporary starting effects from long-run behavior.

Tracking Distributions Over Time

To study long-run behavior, track the probability of being in each state rather than following only one particular run. At the beginning, the state distribution may reflect the initial condition. As the process moves through the MDP, that distribution can change. The early distribution is therefore not automatically the same as the long-run distribution.

continue under πcontinue under πStart AA-heavy distributiondπsame limiting distributionStart BB-heavy distributiondπsame limiting distribution
How do state distributions starting from different initial states change over time, and when do their temporary differences disappear?

The diagram shows the central idea: two processes can begin with different state distributions, yet under the same policy their long-run distributions can agree. The early distributions differ because the initial conditions differ. Their common limiting distribution represents the behavior that remains after those temporary differences have faded.

Ergodicity and the Steady State

In this context, an MDP is ergodic when, under the policy being followed, the long-run state probabilities have well-defined limits and the resulting steady-state distribution does not depend on the initial state. Ergodicity is sufficient to guarantee the existence of the limits used to describe long-run state probabilities.

The steady-state distribution under policy π is written as dπ. For a state s, it represents the limiting probability that the process is in s at time t when actions from earlier times are selected according to π, as t grows without bound. This notation combines three ideas: the particular state s, action selection governed by π, and a long-run limit.

The steady state is preserved by the policy. Once the process is in distribution dπ, continuing to select actions according to π leaves that distribution unchanged. This does not mean that an individual run stops moving. It means that the distribution across possible states remains the same when the policy continues to be applied.

current statecurrent statecurrent statetransition probabilitiestransition probabilitiestransition probabilitieslong-run resultState APolicy πselects actionsdπpreserved distributionState BState C
How does following a fixed policy allow the process to move among states and settle into a long-run distribution?

A Distribution in Motion

Two initial conditions under one policy

Imagine an ergodic MDP with states A, B, and C. Compare two copies of the task that follow the same policy π. The first copy begins with most of its probability on A. The second begins with most of its probability on C.

Initial step: The two state distributions are different because their starting conditions are different. The first distribution is A-heavy, while the second is C-heavy.

Early steps: As the policy and transition probabilities move the processes through the states, each distribution changes. The two distributions may still be noticeably different during this period.

Long-run limit: Because the MDP is being considered as ergodic under policy π, the initial-state influence is temporary. The limiting state distribution is dπ for both copies.

Steady behavior: Once a copy is in dπ, continuing to select actions according to π preserves that distribution, even though individual state visits can continue to vary.

The initial state affects the transient distribution but not the steady-state distribution in this ergodic setting.

apply πapply πcontinuet = 0depends on startt = 1distribution changest = 2temporary differencesremainLong rundπ is preserved
What is the difference between the distribution of states at early time steps and the stable distribution reached in the long run?

The first few time steps describe transient behavior: the distribution is still changing and may still reveal where the process began. The long-run distribution is different conceptually. It is the limiting distribution dπ, determined by the policy and transition probabilities, and it is unchanged when the same policy continues to be applied.

Mistakes About Long-Run Behavior

  • Treating the initial distribution as the steady-state distribution.

    The early distribution can depend strongly on the initial state. In an ergodic MDP, that influence is temporary.

    Fix: Separate the distribution observed at early time steps from the limiting distribution dπ.

  • Thinking that a steady-state distribution means an individual run stops changing states.

    The steady-state property concerns the distribution across states, not the claim that every individual run stops moving.

    Fix: Understand preservation as unchanged state probabilities when actions continue to be selected according to π.

  • Assuming that early actions never matter.

    Early decisions can change the state distribution during the first few steps.

    Fix: Say that their influence is temporary in the ergodic long-run analysis, while the long-run expectation is determined by the policy and transition probabilities.

  • Assuming that changing the policy leaves the same steady-state distribution.

    The long-run state expectation is determined by the policy together with the transition probabilities.

    Fix: Associate the steady-state distribution with the policy under which it was defined.

Check Your Understanding

MEDIUM

An ergodic MDP is run twice under the same policy. Run 1 begins in state A, and Run 2 begins in state B. During the first few steps, their state distributions differ. What should you expect in the long run, and what condition makes that expectation possible?

Hints
  • Separate the early distribution from the limiting distribution.
  • Ask what determines the long-run state expectation.
  • Use the meaning of dπ.

What do you think happens?

Once an ergodic MDP has reached dπ, what happens to the distribution if actions continue to be selected according to π?

  • It is preserved.
  • It must return to the initial distribution.
  • It becomes permanently concentrated in the starting state.
Reveal answer

Answer: It is preserved.

The steady-state distribution is defined so that continuing to apply policy π leaves the distribution unchanged.

Key Takeaways

  1. Ergodicity separates temporary effects of the initial state and early decisions from long-run behavior.
  2. The steady-state distribution dπ is the limiting state distribution under policy π.
  3. The long-run state expectation is determined by the policy and the MDP transition probabilities.
  4. Once the process is in dπ, continuing to select actions according to π preserves the distribution.
  5. Early state distributions can change over time and differ across starting conditions even when their long-run distribution is the same.

Key Takeaways

  • Ergodicity makes the influence of the initial state and early decisions temporary in the long-run analysis.
  • The steady-state distribution dπ is the limiting distribution of states under policy π.
  • This distribution is determined by the policy and transition probabilities and is preserved when π continues to be followed.
  • Transient distributions describe the changing early behavior, while dπ describes the stable long-run behavior.