Stationary Distributions and Markov Transition Matrices
Linear TD(0) relies on the matrix A to control whether weight updates shrink or grow.
The Stability Question
The convergence analysis of linear TD(0) is mainly organized around one matrix: A. The update also contains a vector b, but A determines how the part of the update involving the current weights behaves. The key question is whether A causes weight components to shrink or allows some components to grow.
The central sufficient condition discussed here is that A be positive definite. Under this condition, the linear TD(0) update has the stability behavior needed for the convergence analysis: the current-weight contribution is associated with shrinking rather than amplifying behavior.
Following the Matrix A
For continuing linear TD(0) with γ < 1, A is constructed from three matrix ingredients and the discount factor. Φ is the feature matrix. Its rows contain the feature vector φ(s) for each state. D is a diagonal matrix containing the stationary-distribution values d(s). P is the state-transition matrix under policy π, with entries p(s′ | s).
A = ΦᵀD(I − γP)ΦRead this construction from the inside outward. The transition matrix P describes how the policy moves between states. The factor I − γP combines the identity contribution with discounted transitions. D then weights the rows according to stationary-state occupancy. Finally, Φ on the right and Φᵀ on the left express the resulting state-space structure in the feature representation used by the linear value approximation.
The Inner Matrix
The factor D(I − γP) is the central object in the positive-definiteness argument. It joins two ideas: D applies stationary-distribution weighting, while I − γP represents the identity contribution after accounting for discounted state transitions. The feature matrices then wrap this inner structure to produce A.
The argument examines the row sums and column sums of D(I − γP). Because P is a stochastic matrix and γ < 1, every row sum of this matrix is positive. The remaining part is to establish that the column sums are nonnegative. The stationary-distribution relation d = Pᵀd is used when analyzing those column sums; in the stated continuing, on-policy setting, the resulting components are positive.
1ᵀD(I − γP)The expression 1ᵀD(I − γP) is the row vector containing the column sums of D(I − γP), where 1 is the column vector whose components are all equal to one. Thus, the stationary relation is not an unrelated property of the Markov chain: it is directly used to analyze the column sums needed by the positive-definiteness argument.
Tracing a Constructed A
From Markov Structure to Stability Matrix
Trace the role of each ingredient in A = ΦᵀD(I − γP)Φ without assigning numerical entries.
Start with P: P records the state-transition structure under policy π. Its entries are p(s′ | s), so it describes how probability moves from a current state to a next state.
Apply discounting: The factor γ scales the transition contribution in γP. The inner expression I − γP therefore combines the identity term with discounted transitions.
Apply stationary weighting: Multiplication by D produces D(I − γP). Since D is diagonal and contains d(s), the state-transition structure is weighted by stationary-distribution values.
Return to feature space: The factors Φᵀ and Φ place the inner matrix inside the feature representation, producing A.
Use the stability test: The convergence analysis asks whether the resulting A is positive definite. That condition is the central sufficient condition discussed for the stability analysis.
A is not determined by the feature matrix alone. Its stability role reflects the combined feature, stationary-distribution, discounted-transition, and identity structure.
Reading Shrinkage Correctly
A simple diagonal intuition makes the role of A easier to see. Imagine that A is diagonal. If one diagonal element is negative, the matching diagonal element of I − αA is greater than one. The corresponding weight component is amplified rather than reduced, which can lead to divergence if the process continues.
If every diagonal element is positive, α can be chosen small enough that the diagonal elements of I − αA lie between zero and one. In that simplified picture, the weight components are shrunk by the update. This diagonal case is an intuition for why the sign and definiteness of A matter; the full analysis uses the matrix construction and the positive-definiteness condition rather than checking only isolated diagonal entries.
Avoiding Overclaiming
Treating b as the matrix that controls stability.
The convergence analysis is mainly about A. The vector b appears in the update, but A determines whether the part involving the current weights shrinks or grows.
Fix:
Start by analyzing A and its positive-definiteness condition.Leaving out D when constructing A.
D contains stationary-distribution values and supplies the stationary-state weighting central to the argument.
Fix:
Use the full construction A = ΦᵀD(I − γP)Φ.Treating P as the stationary distribution.
P is the state-transition matrix, while D is the diagonal matrix containing stationary-distribution values.
Fix:
Keep the roles separate: P describes transitions and D supplies stationary-distribution weighting.Claiming that positive definiteness alone is a complete probability-one convergence proof.
Positive definiteness is the central sufficient condition discussed for the stability analysis, but the learning objective explicitly distinguishes it from a complete probability-one convergence proof.
Fix:
State precisely what the condition establishes and acknowledge that a complete probability-one claim requires additional reasoning.
Practice the Construction
Explain, in the correct order, how you would describe the roles of Φ, D, P, and γ when introducing A = ΦᵀD(I − γP)Φ. Then state why D(I − γP), rather than Φ alone, is central to the positive-definiteness argument.
Hints
- Begin with what the rows of Φ contain.
- Separate transition behavior in P from stationary-distribution values in D.
- Mention positive row sums, column-sum analysis, and the relation d = Pᵀd.
- End by stating the exact stability condition involving A.
What do you think happens?
Suppose the current-weight part of a simplified diagonal update contains a diagonal element of I − αA that is greater than one. Does that component shrink or become amplified?
Reveal answer
Answer: It becomes amplified
The diagonal intuition states that a negative diagonal element of A makes the matching diagonal element of I − αA greater than one, so the corresponding weight component is amplified rather than reduced.
Key Takeaways
- A is the matrix that controls whether the current-weight part of the linear TD(0) update shrinks or grows.
- The central sufficient condition in this analysis is that A be positive definite.
- For continuing linear TD(0) with γ < 1, A = ΦᵀD(I − γP)Φ, where Φ contains state feature vectors, D contains stationary-distribution values on its diagonal, and P describes state transitions under the policy.
- D(I − γP) is central because its row sums and column sums support the positive-definiteness argument; the stationary relation d = Pᵀd is used when analyzing the column sums.
- Positive definiteness establishes the relevant stability behavior, but it should not be confused with a complete probability-one convergence proof.
Key Takeaways
- Linear TD(0) convergence analysis centers on the matrix A.
- The matrix is constructed as A = ΦᵀD(I − γP)Φ.
- D(I − γP) combines stationary-state weighting with discounted transition structure and is the key inner matrix in the positive-definiteness argument.
- Positive definiteness supports shrinking update behavior, while a complete probability-one convergence claim requires additional reasoning.