Concepts / Stationary Distributions and Markov Transition Matrices

Stationary Distributions and Markov Transition Matrices

Linear TD(0) relies on the matrix A to control whether weight updates shrink or grow.

  • Programming

The Stability Question

The convergence analysis of linear TD(0) is mainly organized around one matrix: A. The update also contains a vector b, but A determines how the part of the update involving the current weights behaves. The key question is whether A causes weight components to shrink or allows some components to grow.

The central sufficient condition discussed here is that A be positive definite. Under this condition, the linear TD(0) update has the stability behavior needed for the convergence analysis: the current-weight contribution is associated with shrinking rather than amplifying behavior.

positive-definite casenon-shrinking directioncurrent weightsweight vectorshrinking behaviorpositive-definite Agrowing behaviorpossible when a directionis amplified
What changes in the current-weight part of the update when the relevant matrix is positive definite?

Following the Matrix A

For continuing linear TD(0) with γ < 1, A is constructed from three matrix ingredients and the discount factor. Φ is the feature matrix. Its rows contain the feature vector φ(s) for each state. D is a diagonal matrix containing the stationary-distribution values d(s). P is the state-transition matrix under policy π, with entries p(s′ | s).

A = ΦᵀD(I − γP)Φ

Read this construction from the inside outward. The transition matrix P describes how the policy moves between states. The factor I − γP combines the identity contribution with discounted transitions. D then weights the rows according to stationary-state occupancy. Finally, Φ on the right and Φᵀ on the left express the resulting state-space structure in the feature representation used by the linear value approximation.

stationary weightingdiscounted transition termdiscounts Pcombine with feature matricesleft feature factorright feature factorΦᵀfeature representationD(I − γP)weighted discountedtransition structureAstability matrixDstationary-distributionweightsΦfeature representationPstate transitionsγdiscount factor
How do the feature matrix, stationary-distribution matrix, transition matrix, and discount factor combine to form A?

The Inner Matrix

The factor D(I − γP) is the central object in the positive-definiteness argument. It joins two ideas: D applies stationary-distribution weighting, while I − γP represents the identity contribution after accounting for discounted state transitions. The feature matrices then wrap this inner structure to produce A.

The argument examines the row sums and column sums of D(I − γP). Because P is a stochastic matrix and γ < 1, every row sum of this matrix is positive. The remaining part is to establish that the column sums are nonnegative. The stationary-distribution relation d = Pᵀd is used when analyzing those column sums; in the stated continuing, on-policy setting, the resulting components are positive.

1ᵀD(I − γP)

The expression 1ᵀD(I − γP) is the row vector containing the column sums of D(I − γP), where 1 is the column vector whose components are all equal to one. Thus, the stationary relation is not an unrelated property of the Markov chain: it is directly used to analyze the column sums needed by the positive-definiteness argument.

transition structurediscounts transitionsweights statessupportsanalyzed using d = PᵀdPstochastic transitionsD(I − γP)inner matrixpositive row sumsγ < 1discountingpositive column sumsin the stated settingDstationary occupancy
How does D(I − γP) combine stationary-state weighting with discounted transitions in the sign analysis?

Tracing a Constructed A

From Markov Structure to Stability Matrix

Trace the role of each ingredient in A = ΦᵀD(I − γP)Φ without assigning numerical entries.

Start with P: P records the state-transition structure under policy π. Its entries are p(s′ | s), so it describes how probability moves from a current state to a next state.

Apply discounting: The factor γ scales the transition contribution in γP. The inner expression I − γP therefore combines the identity term with discounted transitions.

Apply stationary weighting: Multiplication by D produces D(I − γP). Since D is diagonal and contains d(s), the state-transition structure is weighted by stationary-distribution values.

Return to feature space: The factors Φᵀ and Φ place the inner matrix inside the feature representation, producing A.

Use the stability test: The convergence analysis asks whether the resulting A is positive definite. That condition is the central sufficient condition discussed for the stability analysis.

A is not determined by the feature matrix alone. Its stability role reflects the combined feature, stationary-distribution, discounted-transition, and identity structure.

moves and discounts transitionsweights by stationary occupancyPp(s′ | s)D(I − γP)weighted discountedstructureDdiagonal values d(s)
How does P move probability between states, and how does D represent the long-run occupancy used to weight those transitions?

Reading Shrinkage Correctly

A simple diagonal intuition makes the role of A easier to see. Imagine that A is diagonal. If one diagonal element is negative, the matching diagonal element of I − αA is greater than one. The corresponding weight component is amplified rather than reduced, which can lead to divergence if the process continues.

If every diagonal element is positive, α can be chosen small enough that the diagonal elements of I − αA lie between zero and one. In that simplified picture, the weight components are shrunk by the update. This diagonal case is an intuition for why the sign and definiteness of A matter; the full analysis uses the matrix construction and the positive-definiteness condition rather than checking only isolated diagonal entries.

supportscannot be replaced byA positive definitestability conditionshrinking updatebehaviorcurrent-weight contributioncompleteconvergence proofprobability-one claimadditional reasoningneeded before the fullclaim
What does positive definiteness establish, and what must not be claimed from it alone?

Avoiding Overclaiming

  • Treating b as the matrix that controls stability.

    The convergence analysis is mainly about A. The vector b appears in the update, but A determines whether the part involving the current weights shrinks or grows.

    Fix: Start by analyzing A and its positive-definiteness condition.

  • Leaving out D when constructing A.

    D contains stationary-distribution values and supplies the stationary-state weighting central to the argument.

    Fix: Use the full construction A = ΦᵀD(I − γP)Φ.

  • Treating P as the stationary distribution.

    P is the state-transition matrix, while D is the diagonal matrix containing stationary-distribution values.

    Fix: Keep the roles separate: P describes transitions and D supplies stationary-distribution weighting.

  • Claiming that positive definiteness alone is a complete probability-one convergence proof.

    Positive definiteness is the central sufficient condition discussed for the stability analysis, but the learning objective explicitly distinguishes it from a complete probability-one convergence proof.

    Fix: State precisely what the condition establishes and acknowledge that a complete probability-one claim requires additional reasoning.

Practice the Construction

MEDIUM

Explain, in the correct order, how you would describe the roles of Φ, D, P, and γ when introducing A = ΦᵀD(I − γP)Φ. Then state why D(I − γP), rather than Φ alone, is central to the positive-definiteness argument.

Hints
  • Begin with what the rows of Φ contain.
  • Separate transition behavior in P from stationary-distribution values in D.
  • Mention positive row sums, column-sum analysis, and the relation d = Pᵀd.
  • End by stating the exact stability condition involving A.

What do you think happens?

Suppose the current-weight part of a simplified diagonal update contains a diagonal element of I − αA that is greater than one. Does that component shrink or become amplified?

  • It shrinks
  • It becomes amplified
  • The source information does not distinguish these cases
Reveal answer

Answer: It becomes amplified

The diagonal intuition states that a negative diagonal element of A makes the matching diagonal element of I − αA greater than one, so the corresponding weight component is amplified rather than reduced.

Key Takeaways

  1. A is the matrix that controls whether the current-weight part of the linear TD(0) update shrinks or grows.
  2. The central sufficient condition in this analysis is that A be positive definite.
  3. For continuing linear TD(0) with γ < 1, A = ΦᵀD(I − γP)Φ, where Φ contains state feature vectors, D contains stationary-distribution values on its diagonal, and P describes state transitions under the policy.
  4. D(I − γP) is central because its row sums and column sums support the positive-definiteness argument; the stationary relation d = Pᵀd is used when analyzing the column sums.
  5. Positive definiteness establishes the relevant stability behavior, but it should not be confused with a complete probability-one convergence proof.

Key Takeaways

  • Linear TD(0) convergence analysis centers on the matrix A.
  • The matrix is constructed as A = ΦᵀD(I − γP)Φ.
  • D(I − γP) combines stationary-state weighting with discounted transition structure and is the key inner matrix in the positive-definiteness argument.
  • Positive definiteness supports shrinking update behavior, while a complete probability-one convergence claim requires additional reasoning.