Concepts / Value Prediction in Reinforcement Learning

Value Prediction in Reinforcement Learning

Function approximation connects reinforcement learning value prediction with generalization across states and actions.

  • Programming

Forecasting Future Reward

A reinforcement learning agent often has to make decisions before it knows the complete result of those decisions. Value prediction gives the agent a forecast: how much total reward might be available in the future? A value estimate is therefore a long-run prediction of total future reward, not merely a description of the next immediate result.

The challenge becomes larger when an agent must predict values for many states and actions. Treating every prediction as an entirely separate case may not make good use of experience. Function approximation connects reinforcement learning with generalization: information from some states or actions can support value estimates for other states or actions. The central process is to use a reinforcement learning method to produce value-prediction backups, use those backups as training examples, and then evaluate the resulting approximation.

From Backup to Training Example

A useful way to follow the information is as a three-part chain. First, a reinforcement learning value-prediction method produces backups. Next, those backups serve as training examples for a function approximation method. Finally, the approximation is judged with a performance measure such as MSVE. The direction matters: MSVE evaluates the resulting approximation; it does not produce the training examples.

producesprovidestrainsis evaluated byRL value-predictionmethodBackupTraining exampleFunctionapproximation methodMSVE evaluation
How do a reinforcement learning backup and an existing value estimate become information for improving a value approximation?

The backup is not the final approximation. It is information that can be used as a training example. The approximation method uses such examples to improve its predictions, and MSVE belongs to the later evaluation stage.

Generalization Across States

Function approximation is useful because a reinforcement learning system may need values for many states and actions. An approximation method can connect these predictions instead of treating every one as an entirely separate case. Experience associated with some states or actions can therefore contribute to estimating values for other states or actions.

supportstrainsgeneralizes toExperiencesome states and actionsTraining examplesValue approximationValue estimatesmany states and actions
How can information from experienced states and actions support value estimates for other states and actions?

Evaluating the Approximation

MSVE is a performance measure for function approximation methods. A method may seek to minimize it, so MSVE belongs on the evaluation side of the value-prediction process.

Producing predictions is not enough to establish that an approximation method is performing well. The method is also considered through a measure such as MSVE. In practical reasoning, ask two separate questions: what information trains the approximation, and how is the resulting approximation evaluated? Backups answer the first question by providing training examples. MSVE helps address the second by measuring performance.

Part of the processRole
Reinforcement learning backupProvides information that can become a training example
Function approximation methodUses training examples to represent and improve value predictions
MSVEEvaluates the performance of the approximation method

The roles of backups, approximation, and MSVE

Why Gradient Methods Matter

There is no single universal function approximation method for value prediction. The possible range is large, and the source material notes that too little is known about many alternatives to support a reliable evaluation or recommendation of all of them. For that reason, the discussion focuses on methods based on gradient principles.

Gradient-descent methods are one class of function approximation methods. Linear gradient-descent methods receive particular attention because they are simple, considered promising, and useful for exposing important theoretical issues. Their importance is therefore both practical and explanatory: their relative simplicity makes underlying questions easier to study.

State Values and Action Values

A state value estimate, written as V, assigns an estimate to a state. It describes the total future reward expected from being in that state. An action value estimate, written as Q, assigns an estimate to a specific action. It describes the total future reward expected from taking that particular action.

The distinction is about what is being evaluated. V asks how promising a state is for the agent. Q asks how promising a particular action is. Both are long-run predictions rather than descriptions limited to the next immediate result.

Choosing Between Two Actions

An agent is considering two actions. One action leads to state A, whose value estimate is 8. The other leads to state B, whose value estimate is 3. Which action is preferred when the agent uses state values?

Identify the relevant estimates: The two possible next states have estimates of 8 and 3. These numbers represent predicted total future rewards, not immediate rewards.

Compare the state values: The estimate for state A is larger than the estimate for state B.

Select the preferred decision: Using state values, the agent prefers the action that leads to state A because its relevant state estimate is 8.

The action leading to state A is preferred under the state-value route.

considerconsiderleads toleads tohigher V estimatecomparecomparehigher Q estimatelower Q estimateState-value routeAction 1leads to AState AV = 8Preferred actionAction 2leads to BState BV = 3Action-value routeAction 1Q estimateAction 2Q estimate
How does an agent choose a preferred action when it has state values compared with when it has action values?

With V, the agent compares the values of states reached by available actions. With Q, the agent compares estimates already attached to the actions themselves. In either route, the preferred choice is associated with the larger relevant predicted total future reward.

Mistakes in Value Reasoning

  • Treating a value estimate as an immediate reward

    A value estimate is a long-run prediction of total future reward.

    Fix: Interpret 8 as the predicted total future reward associated with that state or action, depending on whether the estimate is V or Q.

  • Confusing V with Q

    V assigns an estimate to a state, while Q assigns an estimate to a specific action.

    Fix: Ask what is being evaluated: the situation itself or the particular action taken in that situation.

  • Reversing training and evaluation

    Backups provide training examples; MSVE evaluates the performance of the approximation method.

    Fix: Keep the sequence as backup, training example, approximation, then MSVE evaluation.

  • Assuming gradient-based methods are the only possible approximation methods

    Gradient-based methods are a focused subset of a much larger design space.

    Fix: Explain that gradient-based methods receive attention because they are promising and theoretically revealing.

Practice the Decision

EASY

An agent can choose Action X, which leads to a state with value estimate 4, or Action Y, which leads to a state with value estimate 9. If the agent is choosing through state values, which action is preferred? Then describe how the comparison would differ if the agent were using action values instead.

Hints
  • Compare the values of the states reached by the two actions.
  • For action values, compare the estimates attached directly to Action X and Action Y.

What do you think happens?

Which action is preferred through state values: Action X leading to value 4, or Action Y leading to value 9?

  • Action X
  • Action Y
  • There is not enough information to compare the estimates
Reveal answer

Answer: Action Y

Using state values means comparing the values of the states reached by the actions. Since 9 is larger than 4, the action leading to the state valued at 9 is preferred.

Key Takeaways

  1. Function approximation helps reinforcement learning generalize value predictions across many states and actions.
  2. Reinforcement learning backups can become training examples for a function approximation method.
  3. MSVE evaluates the performance of the resulting approximation; it belongs to the evaluation stage.
  4. Gradient-based methods are a focused subset of a much larger design space, with linear gradient-descent methods receiving attention because they are simple, promising, and theoretically useful.
  5. V estimates the long-run reward from a state, while Q estimates the long-run reward from taking a particular action.
  6. Using V means comparing the values of states reached by actions; using Q means comparing the action-value estimates directly.

Key Takeaways

  • Value prediction estimates total future reward rather than only immediate reward.
  • Backups from reinforcement learning methods can supply training examples for function approximation.
  • Function approximation supports generalization across many states and actions, while MSVE evaluates approximation performance.
  • Gradient-based methods receive focused attention because they are promising and theoretically revealing, especially in their linear forms.
  • State values evaluate states; action values evaluate particular actions, and the larger relevant estimate indicates the preferred choice.