Mean Squared Value Error (MSVE)
The on-policy distribution describes states under the target policy π.
Why State Attention Matters
When approximate values are evaluated in reinforcement learning, it matters which states receive attention. The target policy π determines how states are distributed, and the on-policy distribution describes that distribution. In other words, value accuracy is considered in relation to the states that the target policy spends time in.
MSVE is connected to the on-policy distribution: the distribution identifies which states receive emphasis when approximate values are compared with true values.
From Target Policy to State Weights
The on-policy distribution describes states under the target policy π. Following π determines which states are encountered and how much attention each state receives. The resulting distribution supplies the state weighting used when evaluating approximate values.
Counting State Visits
A common way to understand the on-policy distribution is as the fraction of time spent in each state. Imagine observing the target policy over many time steps. Count how often each state is visited, then describe each count as a fraction of the total observation time. States visited more often receive greater emphasis in the distribution.
A Visit-Frequency Interpretation
Suppose an observation under target policy π records visits to several states. How should those observations be interpreted as an on-policy distribution?
Observe: Follow the target policy and record the states encountered over the observation period.
Count: Count how often each state appears in the record.
Convert: Treat each state's count as a fraction of the total number of observed time steps.
Interpret: Use those fractions to identify which states receive more or less attention under the target policy.
The resulting state fractions provide the common visit-frequency interpretation of the on-policy distribution.
Continuing-Task Distribution
In a continuing task, there is no terminal state that ends the task. For this setting, the on-policy distribution is interpreted as the stationary distribution under π. This describes the long-run distribution of states associated with continuing to follow the target policy.
For continuing tasks, replace the intuition of a completed episode with the long-run, stationary distribution of states under the target policy π.
Comparing True and Approximate Values
MSVE evaluates the difference between approximate values and true values across states, while the on-policy distribution determines how much attention each state receives. The comparison is therefore not only about whether values differ; it is also about where those differences occur relative to the states emphasized by the target policy.
Reading the Weighting
Imagine two approximate value systems. In one system, the larger differences from true values occur in states that the target policy spends more time in. In the other, the larger differences occur in states that receive less attention under the target policy. How does the on-policy distribution help interpret the comparison?
Locate the differences: Compare the approximate values with the true values across the relevant states.
Check state emphasis: Use the on-policy distribution to identify how strongly each state is represented under target policy π.
Interpret the overall error: Differences in frequently visited states receive more attention in the on-policy evaluation than differences in less frequently visited states.
The on-policy distribution connects the value comparison to the states that the target policy spends time in.
What RMSVE Tells You
RMSVE gives a rough measure of how much approximate values differ from true values. It is useful as an overall indication of value-estimation discrepancy rather than as a description of one particular state's error.
Common Interpretation Mistakes
Treating the on-policy distribution as independent of the target policy.
The on-policy distribution describes states under the target policy π.
Fix:
Always ask which target policy is being followed before interpreting the state distribution.Ignoring state visitation frequency.
The typical weighting of the on-policy distribution is the fraction of time spent in each state.
Fix:
Interpret the distribution through the relative amount of time that the target policy spends in each state.Using an episodic interpretation for every task.
For continuing tasks, the on-policy distribution is the stationary distribution under π.
Fix:
For a continuing task, use the long-run stationary-distribution interpretation.Reading RMSVE as the error of one particular state.
RMSVE gives a rough measure of how much approximate values differ from true values overall.
Fix:
Use RMSVE as a broad summary, then examine state-level differences and their on-policy emphasis when more detail is needed.
Check Your Understanding
Explain, in your own words, how target policy π, state visitation frequency, the on-policy distribution, MSVE, and RMSVE are connected.
Hints
- Begin with the states encountered when target policy π is followed.
- Relate visit frequency to the fraction of time spent in each state.
- Explain how those state weights affect the comparison between approximate and true values.
- End by describing RMSVE as a rough measure of the resulting value difference.
What do you think happens?
A continuing task is evaluated under target policy π. Should the on-policy distribution be interpreted as a terminal-episode count or as a long-run distribution under π?
Reveal answer
Answer: A long-run stationary distribution under π
For continuing tasks, the on-policy distribution is the stationary distribution under π.
Key Takeaways
- The on-policy distribution describes states under target policy π.
- It is commonly understood as the fraction of time spent in each state.
- In continuing tasks, it is the stationary distribution under π.
- MSVE compares approximate and true values across states while using the on-policy distribution to identify state emphasis.
- RMSVE provides a rough measure of how much approximate values differ from true values.
Key Takeaways
- The target policy π determines the state distribution used for on-policy evaluation.
- The on-policy distribution is commonly described through the fraction of time spent in each state.
- For continuing tasks, that distribution is the stationary distribution under π.
- MSVE connects differences between approximate and true values with the states emphasized by the target policy.
- RMSVE is a rough overall measure of the difference between approximate and true values.