Prediction Objectives in Reinforcement Learning
Value function approximation makes exact prediction for every state impossible in the described setting.
Why Approximation Needs an Objective
Suppose an agent uses a value function approximation instead of storing a separate exact value for every state. In this setting, the learner may not be able to make every state prediction correct at the same time. The estimates are coupled: changing the parameters that improve one state's estimate can also change estimates for other states. An explicit prediction objective is therefore needed to judge the overall quality of the value function.
What do you think happens?
If one shared value-function approximation is updated to improve the estimate for state A, what can happen to the estimates for other states?
Reveal answer
Answer: Other state estimates can also change
Approximation couples state estimates. An update changes shared parameters, so more than one state's prediction can be affected.
Mean Squared Value Error
For a state s, the prediction error is the difference between the approximate value, written as v̂(s, θ), and the true value under policy π, written as vπ(s). Mean Squared Value Error, or MSVE, is the weighted sum of the squared prediction errors across the state space.
MSVE(θ) = Σs d(s)[vπ(s) − v̂(s, θ)]²
The distribution d(s) expresses which state errors matter more to the objective. A state can contribute substantially to MSVE because its prediction is inaccurate, because its weight is large, or because both conditions hold. Changing d(s) changes the priorities of the objective even when the true and approximate values remain unchanged.
Changing the Objective’s Priorities
Two states, two weighting choices
Consider two states. State A has a squared prediction error of 4, and State B has a squared prediction error of 9. Compare the weighted total when d(A) = 0.8 and d(B) = 0.2 with the weighted total when d(A) = 0.2 and d(B) = 0.8.
First weighting choice: With d(A) = 0.8 and d(B) = 0.2, the weighted total is 0.8 × 4 + 0.2 × 9 = 5.0.
Second weighting choice: With d(A) = 0.2 and d(B) = 0.8, the weighted total is 0.2 × 4 + 0.8 × 9 = 8.0.
Interpretation: The state errors have not changed. Only the weights have changed, so the objective now gives more importance to the larger error at State B.
Changing d(s) changes which prediction errors matter most to MSVE.
This example shows why d(s) is part of the prediction objective rather than a decorative detail. If the learner's priorities emphasize one group of states, errors in those states contribute more strongly to the total measure. The objective is therefore a statement about both prediction accuracy and which states deserve attention.
From Targets to Policy Evaluation
A prediction algorithm in reinforcement learning estimates a quantity whose value depends on how features of the environment are expected to unfold in the future. The target may be future reward, an upcoming stimulus, or another numerical-valued environmental feature.
A useful way to follow a prediction task is to identify the target quantity, consider how that quantity is expected to unfold as the agent interacts with the environment, and then produce an estimate for each relevant state. When the target is future reward expected under a policy, the prediction algorithm is performing policy evaluation.
Reward and Other Environmental Features
| Prediction target | What is estimated | Relationship to policy evaluation |
|---|---|---|
| Future reward | Expected reward as the agent continues interacting with the environment | Estimating future reward under a policy is policy evaluation |
| Upcoming stimulus | Expected future value of a stimulus or other numerical environmental feature | It is a prediction task, but it is not automatically policy evaluation |
| Other numerical environmental feature | Expected future behavior of the selected feature | It uses the broader prediction idea without requiring reward as the target |
From Evaluation to Improvement
Policy evaluation contributes to, but is not identical to, policy improvement. A prediction algorithm estimates what future reward is associated with the policy being evaluated. That information can support a broader process in which the agent seeks to improve its policy. The prediction is therefore useful as an input to improvement, not a complete description of improvement itself.
Classical Conditioning Boundary
Prediction algorithms and classical conditioning overlap because both can involve learning expectations about upcoming stimuli. The connection is limited, however. The stimulus being predicted does not have to be a reward or penalty that evaluates an earlier action. Therefore, a prediction task can share the idea of anticipating an upcoming stimulus without being a classical-conditioning task.
Treating every prediction task as policy evaluation
Policy evaluation specifically concerns estimating future reward under a policy.
Fix:
Use the broader term prediction for numerical environmental targets, and use policy evaluation when the target is future reward under a policy.Treating policy evaluation and policy improvement as the same operation
Evaluation provides information that can contribute to improvement, but the two concepts are not identical.
Fix:
Describe evaluation as estimating the outcome associated with a policy and improvement as the broader process of changing or selecting policy behavior.Assuming an update affects only the state being considered
Coupled state estimates mean that changing shared approximation parameters can change several state predictions.
Fix:
Judge the total prediction quality across states with an explicit objective.Ignoring d(s) when interpreting MSVE
MSVE weights squared errors by d(s), so a smaller error with a larger weight can matter more.
Fix:
Consider both the squared prediction error and the state's weighting.
Check Your Understanding
Explain why an explicit objective is needed when a value function is approximated. Then write the MSVE expression and describe what would change if d(s) placed more weight on a state with a large prediction error. Finally, distinguish future reward prediction from prediction of another numerical environmental feature.
Hints
- Mention that state estimates are coupled by shared approximation parameters.
- Use the structure MSVE(θ) = Σs d(s)[vπ(s) − v̂(s, θ)]².
- Explain that increasing a state's weight increases the importance of its squared error.
- State that future reward prediction under a policy is policy evaluation, while other numerical targets remain broader prediction tasks.
Key Takeaways
- Value function approximation couples state estimates, so improving one estimate can change others and create or alter errors in multiple states.
- MSVE measures overall prediction quality as the weighted sum of squared differences between approximate and true state values.
- The state weighting distribution d(s) determines which prediction errors matter most to the objective.
- Prediction algorithms estimate future-dependent numerical quantities; predicting future reward under a policy is policy evaluation.
- Prediction supports policy improvement but is not identical to improvement, and its connection to classical conditioning is limited to the shared idea of anticipating upcoming stimuli.