Concepts / True and Approximate Value Functions

True and Approximate Value Functions

Value function approximation makes exact prediction for every state impossible in the described setting.

  • Programming

When Exact Prediction Is Impossible

A true value function assigns a true value to each state under a policy. An approximate value function tries to predict those values, but in the setting described here it cannot represent every state perfectly. That creates a problem: checking whether each individual state is exactly correct is no longer enough to judge the quality of the whole prediction.

Approximation couples state predictions. The same parameters used by an approximate value function can influence several state estimates. As a result, improving one estimate may change other estimates as well. Since all states may not be made exact at the same time, the learner needs an explicit objective that says how the total prediction quality should be evaluated.

comparecomparecompareState s1true valueState s1approximate valueState s2true valueState s2approximate valueState s3true valueState s3approximate value
What does it mean for an approximate value function to match or differ from the true value function at multiple states?

A Coupled Update

What do you think happens?

Suppose an update improves the approximate estimate for one state. What can happen to the estimates for other states when those estimates are coupled?

  • Only the updated state can change
  • Other state estimates can also change
  • All state estimates must become exact
Reveal answer

Answer: Other state estimates can also change

The source describes approximate predictions as coupled: changing the parameters used by the value function can change several state estimates. Therefore, an update can improve one state while making another state more or less accurate.

affected by shared parametersaffected by shared parametersaffected by shared parameterschanges estimatechanges estimatechanges estimateState s1approximate valueState s1new approximate valueState s2approximate valueParameter updateshared approximationState s2new approximate valueState s3approximate valueState s3new approximate value
When one approximate value estimate is updated, which states can become more or less accurate, and how can their errors change?

The important consequence is that an update does not have to produce one isolated error change. One state may become more accurate, while another becomes less accurate or remains inaccurate. The prediction method therefore needs a way to combine errors across the state space.

The State Weighting Distribution

The state weighting distribution d(s) expresses which state errors matter more when prediction quality is evaluated. It gives each state a role in the overall objective, so the same set of approximate values can receive different evaluations when the state weights change.

assigned by d(s)assigned by d(s)assigned by d(s)State s1d(s1)State weightimportance for s1State s2d(s2)State weightimportance for s2State s3d(s3)State weightimportance for s3
How does d(s) assign different importance or frequency weights to states when evaluating prediction quality?

Same errors, different priorities

Consider three states whose prediction errors are already known. Compare the evaluation when one state is given a larger weight.

Identify the errors: For each state, compare the approximate value with the true value. The resulting differences are the per-state prediction errors.

Apply the state weights: Each squared error is multiplied by that state's d(s). A larger d(s) makes that state's error contribute more strongly to the total objective.

Change one priority: If the weight for state s2 increases while the approximate and true values stay the same, the contribution from s2 increases relative to the contributions from the other states.

Changing d(s) changes which errors matter most, even when the predictions and true values do not change.

Mean Squared Value Error

For a state s, the prediction error is the difference between the approximate value, written as ˆv(s, θ), and the true value under policy π, written as vπ(s). Mean Squared Value Error, or MSVE, is the weighted sum of the squared differences between these approximate and true state values.

MSVE(θ) = Σs d(s) [ˆv(s, θ) − vπ(s)]²

subtractsquaremultiply by d(s)add over statesApproximate and truevaluesValue differenceˆv(s, θ) − vπ(s)Squared error[ˆv(s, θ) − vπ(s)]²Weighted errord(s) times squared errorMSVEsum over states
How are per-state value errors squared, weighted by d(s), and combined into one overall error measure?

MSVE turns many state-level comparisons into one objective. It does not ask only whether one state is correct. Instead, it combines every state's squared prediction error after applying the state weighting distribution. The arithmetic is less important than the structure: error, square, weight, and sum.

Reading the Objective

Why a large error may not dominate

Suppose two states have different prediction errors. State s1 has a larger squared error, but state s2 has a larger d(s). What determines each state's contribution to MSVE?

Evaluate state s1: Square the difference between ˆv(s1, θ) and vπ(s1), then multiply that squared error by d(s1).

Evaluate state s2: Square the difference between ˆv(s2, θ) and vπ(s2), then multiply that squared error by d(s2).

Compare contributions: The larger contribution is determined by the product of the squared error and the state's weight, not by the error alone.

A state affects MSVE through both its prediction accuracy and its state weight.

weighted contributionbefore weighting changeafter weighting changeState s1squared errord(s1)lower weightd(s1)same weightState s2squared errord(s2)lower weightd(s2)higher weight
How does increasing the weight of one state change its contribution to the overall prediction error compared with other states?

If the approximate and true values remain fixed, changing d(s) can still change MSVE. Increasing the weight of one state makes that state's squared error more important in the combined measure. This is why the prediction objective must specify not only what counts as error, but also which state errors receive greater attention.

Common Reasoning Mistakes

  • Assuming that improving one state leaves every other state unchanged.

    Approximate state predictions are coupled, so changing the parameters can change several state estimates.

    Fix: Track the possible effects of an update across multiple states.

  • Evaluating approximation using only the largest individual error.

    MSVE combines weighted squared errors over the state space.

    Fix: Consider every state's squared error together with its d(s) weight.

  • Treating d(s) as irrelevant bookkeeping.

    The state weighting distribution expresses which state errors matter more.

    Fix: Include d(s) when interpreting each state's contribution to the objective.

  • Thinking that a lower error always means a lower contribution to MSVE.

    A larger state weight can make a smaller squared error contribute more strongly.

    Fix: Compare weighted squared errors rather than errors alone.

Practice

MEDIUM

Explain in your own words why a value function approximation method needs an explicit prediction objective. Then describe what happens to the objective if the approximate and true values stay fixed but d(s) is increased for one state.

Hints
  • Mention that approximate predictions can be coupled across states.
  • Use the structure of MSVE: difference, square, weight, and sum.
  • Focus on how the selected state contributes more strongly after its weight increases.

Key Takeaways

  1. Approximation can make exact prediction for every state impossible.
  2. Because state estimates are coupled, improving one estimate can change predictions and errors in other states.
  3. The state weighting distribution d(s) expresses which state errors matter more.
  4. MSVE is the sum over states of d(s) multiplied by the squared difference between ˆv(s, θ) and vπ(s).
  5. Changing state weights changes the priorities of the prediction objective even when the predictions themselves stay the same.

Key Takeaways

  • Approximate value functions may not be able to match the true value at every state.
  • Coupled predictions mean that one update can affect errors in multiple states.
  • The distribution d(s) determines how strongly each state's error matters.
  • MSVE combines squared state errors after weighting each one by d(s).
  • A prediction objective makes approximation trade-offs explicit.