Concepts / Policies and Transition Probabilities

Policies and Transition Probabilities

Ergodicity separates temporary starting effects from long-run behavior.

  • Programming

From Starting Point to Long Run

Imagine observing a continuing task immediately after it begins. At the first few time steps, the states you observe can depend strongly on where the MDP started and on the actions selected early by the agent. The central question is what remains relevant after the process has continued for a long time. Ergodicity gives a precise answer: the influence of the starting state and early decisions is temporary, while long-run state behavior is determined by the policy and the MDP transition probabilities.

continued transitionscontinued transitionslong-run evolutionlong-run evolutionInitialdistribution Aearly timeDistribution A laterunder policy πSteady-statedistributionlong-run behaviorInitialdistribution Bearly timeDistribution B laterunder policy π
How do state distributions that begin differently gradually become similar under the same policy?

How a Policy Shapes Transitions

A policy determines how actions are selected as the process visits states. The MDP also has transition probabilities describing how the process moves between states. Together, the policy and the transition probabilities determine the long-run expectation of being in a state. This is why the long-run behavior cannot be understood from the policy alone or from the starting state alone.

current contextselectsused withproducesCurrent statestate at time tSelected actionaccording to πNext-statedistributionstate at time t plus 1Policy πaction selectionTransitionprobabilitiesstate movement
How does choosing actions according to a policy determine the probabilities of moving from one state to another?

The long-run state expectation is determined by two ingredients: the policy used to select actions and the MDP transition probabilities.

The Early Distribution

To reason about ergodicity, track the probability of being in each state rather than following only one particular run. At early times, this distribution may change as the process moves away from its initial condition. The first distribution reflects where the MDP began, and early distributions can also reflect the actions selected during those first steps. These observations describe the transient part of the process: behavior that matters initially but does not determine the long-run expectation in an ergodic MDP.

process continuesprocess continueseffects fadeInitial distributiontime 0Changing distributionearly stepChanging distributionlater early stepSteady-statedistributionlong-run limit
How does the distribution over states change during the first few steps, and how is that different from the eventual steady-state distribution?

Reading the Steady-State Distribution

The steady-state distribution dπ is the limiting distribution of states under policy π. For a state s, it represents the long-run probability that the process is in s when actions from the earlier time steps are selected according to π.

The notation dπ combines three ideas. First, the question concerns a particular state s. Second, the actions leading up to the observation are selected according to policy π. Third, the process is considered as time grows without bound. Ergodicity is sufficient to guarantee the existence of the limits used to describe these long-run state probabilities. The steady-state distribution is also assumed to exist independently of the initial state S0.

part of distributionpart of distributionpart of distributionState Adπ(A)Same distributionafter applying πState Bdπ(B)State Cdπ(C)
What does the long-run probability of being in each state look like once the effects of the starting state have faded?

Once the process is in the steady-state distribution, selecting actions according to policy π preserves that distribution. In other words, continuing to apply the policy leaves the state distribution unchanged.

A Distribution-Level Walkthrough

Separating Temporary and Long-Run Effects

Two versions of the same continuing MDP begin from different initial state distributions. Both then select actions according to the same policy π. What should you compare at early times and what should you compare after the process has continued for a long time?

Step 1: Compare the beginning: At the first few time steps, the two state distributions may differ because the MDPs began differently. Early action selections can also influence the distributions observed at these steps.

Step 2: Follow the same policy: As the process continues, both versions use policy π together with the MDP transition probabilities. These are the ingredients that determine long-run state behavior.

Step 3: Identify the long-run object: The relevant long-run object is the steady-state distribution dπ, defined as the limiting state distribution under policy π.

Step 4: Remove the wrong dependency: Under ergodicity, the steady-state distribution is assumed to exist independently of the initial state S0. Therefore, the different starting distributions explain temporary differences, not different long-run steady-state distributions.

Step 5: Check preservation: Once either process is in the steady-state distribution, continuing to select actions according to π preserves that distribution.

The initial state and early decisions can affect what is observed first, but the long-run state expectation is determined by policy π and the MDP transition probabilities. The steady-state distribution is the limiting distribution and is preserved when the policy continues to be applied.

temporary influence fadestemporary influence fadeslong-run behaviorInitial stateaffects early observationsPolicy πcontinued action selectionEarly decisionsaffect early distributionSteady-statedistributionindependent of S0
Which parts of the trajectory are affected by the initial state and early actions, and when do those effects disappear from long-run behavior?

Mistakes About Ergodicity

  • Treating the initial state as the long-run state distribution.

    The initial state can strongly affect the first few time steps, but ergodicity separates this temporary effect from long-run behavior.

    Fix: Track the limiting distribution under the policy and transition probabilities rather than carrying the initial condition forward indefinitely.

  • Calling the distribution at the first few time steps the steady-state distribution.

    The early distribution may still be changing as the process moves away from its initial condition.

    Fix: Reserve the term steady-state distribution for the limiting distribution under policy π.

  • Assuming that the steady-state distribution is just the result of the policy.

    The long-run state expectation is determined by both the policy and the transition probabilities.

    Fix: Always interpret dπ as a property of the policy together with the MDP dynamics.

  • Thinking that a steady-state distribution must continue changing after it is reached.

    The defining preservation property says that applying policy π in the steady-state distribution leaves the distribution unchanged.

    Fix: Understand steady state as a distribution preserved by continued action selection according to π.

Check Your Understanding

MEDIUM

A continuing MDP is observed at time 0 and again after a long period. At time 0, its state distribution depends strongly on the initial state. Later, actions continue to be selected according to policy π. Explain which observation is more likely to reflect the steady-state distribution dπ, and identify which information can have only a temporary influence.

Hints
  • Look for the observation taken after the process has continued for a long time.
  • Separate the initial state and early decisions from the policy and transition probabilities.
  • Use the preservation property of the steady-state distribution.

What do you think happens?

If two processes use the same policy and MDP transition probabilities but begin from different initial states, what should ergodicity lead you to expect about their long-run state distributions?

Reveal answer

Answer: Their initial differences are temporary, and the long-run behavior is described by the same steady-state distribution under the policy.

Ergodicity separates the influence of the starting state from long-run behavior. The steady-state distribution is assumed to exist independently of the initial state, while the long-run expectation is determined by the policy and transition probabilities.

Long-Run Interpretation

  1. Ergodicity separates temporary effects from long-run behavior in an MDP.
  2. The initial state and early decisions can strongly affect the state distribution during the first few time steps.
  3. The steady-state distribution dπ is the limiting state distribution under policy π.
  4. The long-run state expectation is determined by the policy and the MDP transition probabilities.
  5. Once the process is in the steady-state distribution, continuing to select actions according to π preserves that distribution.

Key Takeaways

  • Ergodicity makes the influence of the starting state and early decisions temporary.
  • State distributions can change during the early steps as the process moves away from its initial condition.
  • The steady-state distribution dπ is the limiting distribution under policy π.
  • The long-run distribution depends on the policy and transition probabilities rather than on the initial state.
  • A steady-state distribution is preserved when actions continue to be selected according to the same policy.