Defining the Prediction Objective in Reinforcement Learning
The on-policy distribution describes states under the target policy π.
Why State Attention Matters
A prediction objective in reinforcement learning does not treat every state as equally important by default. The relevant question is which states receive attention under the target policy π. The on-policy distribution describes how states are distributed under that policy, connecting value-prediction evaluation to the states that the target policy spends time in.
The target policy determines the state distribution used to describe on-policy prediction.
From Policy to Visited States
Start with the target policy π. The on-policy distribution is the distribution of states under that policy: it identifies which states the target policy visits and how much attention those states receive. In practical terms, this distribution is commonly described by the fraction of time spent in each state.
Reading Time Fractions
The phrase fraction of time gives an intuitive interpretation of the on-policy distribution. Imagine observing states while the target policy is being followed. A state that appears often receives a larger share of the distribution; a state that appears less often receives a smaller share. This is why the prediction objective is connected to the states on which the target policy spends time.
Turning Visits into State Emphasis
Consider the generated observation sequence State A, State B, State A, State A, State B under the target policy π. What does this sequence suggest about the on-policy distribution?
Count appearances: State A appears three times, while State B appears two times.
Interpret the counts: The sequence spends more observed time in State A than in State B.
Connect to prediction: The on-policy distribution therefore places greater emphasis on State A than on State B in this illustration.
The on-policy distribution is interpreted through the relative time spent in each state under the target policy.
Continuing-Task Interpretation
In a continuing task, the process continues rather than being organized around an ending terminal state. In this setting, the on-policy distribution is interpreted as the stationary distribution under the target policy π. The key idea remains the same: the distribution describes the states associated with following π, but for a continuing task it is characterized as the stationary distribution.
Interpreting RMSVE
RMSVE gives a rough measure of how much approximate values differ from true values.
When evaluating an approximate value function, compare its predicted values with the true values across states. RMSVE summarizes the size of that difference in a rough overall measure. In the prediction objective, the on-policy distribution supplies the emphasis: differences in states that receive more attention under π contribute more strongly to the evaluation than differences in states receiving less attention.
Using RMSVE as an Evaluation Signal
Suppose an approximate value function is evaluated across states emphasized by the target policy π. What does a larger RMSVE indicate?
Compare values: Examine the approximate values and the true values for the states being evaluated.
Account for state emphasis: Interpret the differences with attention to the on-policy distribution, which describes how states are distributed under π.
Read the measure: A larger RMSVE indicates a larger rough overall difference between the approximate and true values.
RMSVE is an overall rough indicator of the difference between approximate and true values, interpreted with the state emphasis supplied by the on-policy distribution.
Weighting the Prediction Objective
The prediction objective is weighted by the on-policy distribution. This means the objective emphasizes errors according to how states are distributed under the target policy π. The distribution is therefore not an unrelated statistic: it determines which state-wise differences between approximate and true values receive greater attention when values are evaluated.
To understand what a value-prediction objective emphasizes, first identify the state distribution induced by the target policy.
Common Interpretation Errors
Treating the on-policy distribution as independent of the target policy.
The on-policy distribution describes states under the target policy π.
Fix:
Always connect the distribution to the states visited under the target policy.Forgetting the time-spent interpretation.
The typical weighting is the fraction of time spent in each state.
Fix:
Ask how much time the target policy spends in each state.Applying only an ending-task interpretation to a continuing task.
For continuing tasks, the on-policy distribution is the stationary distribution under π.
Fix:
Use the stationary-distribution interpretation for continuing tasks.Treating RMSVE as an exact explanation of why values are wrong.
RMSVE is described as a rough measure of how much approximate values differ from true values.
Fix:
Interpret RMSVE as an overall rough difference measure, together with the state emphasis supplied by the on-policy distribution.
Check Your Understanding
A target policy π spends most of its time in State A and less time in State B. Explain how this affects the on-policy distribution and the weighting of value-prediction differences. Then state how the interpretation changes when the task is continuing.
Hints
- Begin with the phrase states under the target policy π.
- Use fraction of time spent in each state to explain the weighting.
- For the continuing case, use the term stationary distribution under π.
What do you think happens?
If the approximate values differ more from the true values across states emphasized by the target policy, should RMSVE indicate a larger or smaller rough difference?
Reveal answer
Answer: A larger rough difference
RMSVE gives a rough measure of how much approximate values differ from true values, while the on-policy distribution identifies the state emphasis used in the evaluation.
Key Takeaways
- The on-policy distribution describes states under the target policy π.
- It is commonly understood as the fraction of time spent in each state.
- For continuing tasks, it is the stationary distribution under π.
- The distribution determines which states receive emphasis in the prediction objective.
- RMSVE gives a rough measure of how much approximate values differ from true values.
Key Takeaways
- The target policy π defines the state distribution relevant to on-policy prediction.
- The on-policy distribution is commonly interpreted through the fraction of time spent in each state.
- In continuing tasks, the on-policy distribution is the stationary distribution under π.
- The on-policy distribution weights the prediction objective toward the states emphasized by π.
- RMSVE provides a rough overall measure of the difference between approximate and true values.