Concepts / Prediction Objectives in Reinforcement Learning

Prediction Objectives in Reinforcement Learning

Value function approximation makes exact prediction for every state impossible in the described setting.

  • Programming

Why Approximation Needs an Objective

Suppose an agent uses a value function approximation instead of storing a separate exact value for every state. In this setting, the learner may not be able to make every state prediction correct at the same time. The estimates are coupled: changing the parameters that improve one state's estimate can also change estimates for other states. An explicit prediction objective is therefore needed to judge the overall quality of the value function.

What do you think happens?

If one shared value-function approximation is updated to improve the estimate for state A, what can happen to the estimates for other states?

  • Only state A can change
  • Other state estimates can also change
  • All state estimates must become exact
  • No state estimate can change
Reveal answer

Answer: Other state estimates can also change

Approximation couples state estimates. An update changes shared parameters, so more than one state's prediction can be affected.

shared-parameter updateshared-parameter updateshared-parameter updateState Aestimate 4State Aestimate 5State Bestimate 6State Bestimate 5State Cestimate 3State Cestimate 3
When shared approximation parameters are updated for one state, which state predictions can change and where can new errors appear?

Mean Squared Value Error

For a state s, the prediction error is the difference between the approximate value, written as v̂(s, θ), and the true value under policy π, written as vπ(s). Mean Squared Value Error, or MSVE, is the weighted sum of the squared prediction errors across the state space.

MSVE(θ) = Σs d(s)[vπ(s) − v̂(s, θ)]²

The distribution d(s) expresses which state errors matter more to the objective. A state can contribute substantially to MSVE because its prediction is inaccurate, because its weight is large, or because both conditions hold. Changing d(s) changes the priorities of the objective even when the true and approximate values remain unchanged.

weighted bymultiplyweighted bymultiplyState Asquared error 4d(A)0.8State A contribution3.2State Bsquared error 9d(B)0.2State B contribution1.8
How are each state's squared prediction error and its weight d(s) combined, and what changes when the weights change?

Changing the Objective’s Priorities

Two states, two weighting choices

Consider two states. State A has a squared prediction error of 4, and State B has a squared prediction error of 9. Compare the weighted total when d(A) = 0.8 and d(B) = 0.2 with the weighted total when d(A) = 0.2 and d(B) = 0.8.

First weighting choice: With d(A) = 0.8 and d(B) = 0.2, the weighted total is 0.8 × 4 + 0.2 × 9 = 5.0.

Second weighting choice: With d(A) = 0.2 and d(B) = 0.8, the weighted total is 0.2 × 4 + 0.8 × 9 = 8.0.

Interpretation: The state errors have not changed. Only the weights have changed, so the objective now gives more importance to the larger error at State B.

Changing d(s) changes which prediction errors matter most to MSVE.

This example shows why d(s) is part of the prediction objective rather than a decorative detail. If the learner's priorities emphasize one group of states, errors in those states contribute more strongly to the total measure. The objective is therefore a statement about both prediction accuracy and which states deserve attention.

From Targets to Policy Evaluation

A prediction algorithm in reinforcement learning estimates a quantity whose value depends on how features of the environment are expected to unfold in the future. The target may be future reward, an upcoming stimulus, or another numerical-valued environmental feature.

A useful way to follow a prediction task is to identify the target quantity, consider how that quantity is expected to unfold as the agent interacts with the environment, and then produce an estimate for each relevant state. When the target is future reward expected under a policy, the prediction algorithm is performing policy evaluation.

determines behaviorproduces future quantityestimate by statePolicy πagent's behaviorEnvironmentinteractionfuture unfoldsFuture rewardprediction targetState predictionvalue under π
How does a prediction algorithm take a policy and environmental outcomes and produce predictions for states?

Reward and Other Environmental Features

Prediction targetWhat is estimatedRelationship to policy evaluation
Future rewardExpected reward as the agent continues interacting with the environmentEstimating future reward under a policy is policy evaluation
Upcoming stimulusExpected future value of a stimulus or other numerical environmental featureIt is a prediction task, but it is not automatically policy evaluation
Other numerical environmental featureExpected future behavior of the selected featureIt uses the broader prediction idea without requiring reward as the target
target choicetarget choiceestimateestimateExpected futurecommon prediction basisFuture rewardpolicy evaluationValue under πreward predictionEnvironmental featurenumerical targetFeature predictionnon-reward prediction
What is shared, and what differs, when predicting future reward versus another numerical environmental feature?

From Evaluation to Improvement

Policy evaluation contributes to, but is not identical to, policy improvement. A prediction algorithm estimates what future reward is associated with the policy being evaluated. That information can support a broader process in which the agent seeks to improve its policy. The prediction is therefore useful as an input to improvement, not a complete description of improvement itself.

evaluateprovide predictionssupport improvementevaluate againPolicy πcurrent behaviorPolicy evaluationpredict future rewardAction comparisonuse predictionsImproved policybroader improvement process
How can predictions about future outcomes feed into choosing better actions and iteratively improving a policy?

Classical Conditioning Boundary

Prediction algorithms and classical conditioning overlap because both can involve learning expectations about upcoming stimuli. The connection is limited, however. The stimulus being predicted does not have to be a reward or penalty that evaluates an earlier action. Therefore, a prediction task can share the idea of anticipating an upcoming stimulus without being a classical-conditioning task.

  • Treating every prediction task as policy evaluation

    Policy evaluation specifically concerns estimating future reward under a policy.

    Fix: Use the broader term prediction for numerical environmental targets, and use policy evaluation when the target is future reward under a policy.

  • Treating policy evaluation and policy improvement as the same operation

    Evaluation provides information that can contribute to improvement, but the two concepts are not identical.

    Fix: Describe evaluation as estimating the outcome associated with a policy and improvement as the broader process of changing or selecting policy behavior.

  • Assuming an update affects only the state being considered

    Coupled state estimates mean that changing shared approximation parameters can change several state predictions.

    Fix: Judge the total prediction quality across states with an explicit objective.

  • Ignoring d(s) when interpreting MSVE

    MSVE weights squared errors by d(s), so a smaller error with a larger weight can matter more.

    Fix: Consider both the squared prediction error and the state's weighting.

Check Your Understanding

MEDIUM

Explain why an explicit objective is needed when a value function is approximated. Then write the MSVE expression and describe what would change if d(s) placed more weight on a state with a large prediction error. Finally, distinguish future reward prediction from prediction of another numerical environmental feature.

Hints
  • Mention that state estimates are coupled by shared approximation parameters.
  • Use the structure MSVE(θ) = Σs d(s)[vπ(s) − v̂(s, θ)]².
  • Explain that increasing a state's weight increases the importance of its squared error.
  • State that future reward prediction under a policy is policy evaluation, while other numerical targets remain broader prediction tasks.

Key Takeaways

  • Value function approximation couples state estimates, so improving one estimate can change others and create or alter errors in multiple states.
  • MSVE measures overall prediction quality as the weighted sum of squared differences between approximate and true state values.
  • The state weighting distribution d(s) determines which prediction errors matter most to the objective.
  • Prediction algorithms estimate future-dependent numerical quantities; predicting future reward under a policy is policy evaluation.
  • Prediction supports policy improvement but is not identical to improvement, and its connection to classical conditioning is limited to the shared idea of anticipating upcoming stimuli.