Value Prediction in Reinforcement Learning
Function approximation connects reinforcement learning value prediction with generalization across states and actions.
Forecasting Future Reward
A reinforcement learning agent often has to make decisions before it knows the complete result of those decisions. Value prediction gives the agent a forecast: how much total reward might be available in the future? A value estimate is therefore a long-run prediction of total future reward, not merely a description of the next immediate result.
The challenge becomes larger when an agent must predict values for many states and actions. Treating every prediction as an entirely separate case may not make good use of experience. Function approximation connects reinforcement learning with generalization: information from some states or actions can support value estimates for other states or actions. The central process is to use a reinforcement learning method to produce value-prediction backups, use those backups as training examples, and then evaluate the resulting approximation.
From Backup to Training Example
A useful way to follow the information is as a three-part chain. First, a reinforcement learning value-prediction method produces backups. Next, those backups serve as training examples for a function approximation method. Finally, the approximation is judged with a performance measure such as MSVE. The direction matters: MSVE evaluates the resulting approximation; it does not produce the training examples.
The backup is not the final approximation. It is information that can be used as a training example. The approximation method uses such examples to improve its predictions, and MSVE belongs to the later evaluation stage.
Generalization Across States
Function approximation is useful because a reinforcement learning system may need values for many states and actions. An approximation method can connect these predictions instead of treating every one as an entirely separate case. Experience associated with some states or actions can therefore contribute to estimating values for other states or actions.
Evaluating the Approximation
MSVE is a performance measure for function approximation methods. A method may seek to minimize it, so MSVE belongs on the evaluation side of the value-prediction process.
Producing predictions is not enough to establish that an approximation method is performing well. The method is also considered through a measure such as MSVE. In practical reasoning, ask two separate questions: what information trains the approximation, and how is the resulting approximation evaluated? Backups answer the first question by providing training examples. MSVE helps address the second by measuring performance.
| Part of the process | Role |
|---|---|
| Reinforcement learning backup | Provides information that can become a training example |
| Function approximation method | Uses training examples to represent and improve value predictions |
| MSVE | Evaluates the performance of the approximation method |
The roles of backups, approximation, and MSVE
Why Gradient Methods Matter
There is no single universal function approximation method for value prediction. The possible range is large, and the source material notes that too little is known about many alternatives to support a reliable evaluation or recommendation of all of them. For that reason, the discussion focuses on methods based on gradient principles.
Gradient-descent methods are one class of function approximation methods. Linear gradient-descent methods receive particular attention because they are simple, considered promising, and useful for exposing important theoretical issues. Their importance is therefore both practical and explanatory: their relative simplicity makes underlying questions easier to study.
State Values and Action Values
A state value estimate, written as V, assigns an estimate to a state. It describes the total future reward expected from being in that state. An action value estimate, written as Q, assigns an estimate to a specific action. It describes the total future reward expected from taking that particular action.
The distinction is about what is being evaluated. V asks how promising a state is for the agent. Q asks how promising a particular action is. Both are long-run predictions rather than descriptions limited to the next immediate result.
Choosing Between Two Actions
An agent is considering two actions. One action leads to state A, whose value estimate is 8. The other leads to state B, whose value estimate is 3. Which action is preferred when the agent uses state values?
Identify the relevant estimates: The two possible next states have estimates of 8 and 3. These numbers represent predicted total future rewards, not immediate rewards.
Compare the state values: The estimate for state A is larger than the estimate for state B.
Select the preferred decision: Using state values, the agent prefers the action that leads to state A because its relevant state estimate is 8.
The action leading to state A is preferred under the state-value route.
With V, the agent compares the values of states reached by available actions. With Q, the agent compares estimates already attached to the actions themselves. In either route, the preferred choice is associated with the larger relevant predicted total future reward.
Mistakes in Value Reasoning
Treating a value estimate as an immediate reward
A value estimate is a long-run prediction of total future reward.
Fix:
Interpret 8 as the predicted total future reward associated with that state or action, depending on whether the estimate is V or Q.Confusing V with Q
V assigns an estimate to a state, while Q assigns an estimate to a specific action.
Fix:
Ask what is being evaluated: the situation itself or the particular action taken in that situation.Reversing training and evaluation
Backups provide training examples; MSVE evaluates the performance of the approximation method.
Fix:
Keep the sequence as backup, training example, approximation, then MSVE evaluation.Assuming gradient-based methods are the only possible approximation methods
Gradient-based methods are a focused subset of a much larger design space.
Fix:
Explain that gradient-based methods receive attention because they are promising and theoretically revealing.
Practice the Decision
An agent can choose Action X, which leads to a state with value estimate 4, or Action Y, which leads to a state with value estimate 9. If the agent is choosing through state values, which action is preferred? Then describe how the comparison would differ if the agent were using action values instead.
Hints
- Compare the values of the states reached by the two actions.
- For action values, compare the estimates attached directly to Action X and Action Y.
What do you think happens?
Which action is preferred through state values: Action X leading to value 4, or Action Y leading to value 9?
Reveal answer
Answer: Action Y
Using state values means comparing the values of the states reached by the actions. Since 9 is larger than 4, the action leading to the state valued at 9 is preferred.
Key Takeaways
- Function approximation helps reinforcement learning generalize value predictions across many states and actions.
- Reinforcement learning backups can become training examples for a function approximation method.
- MSVE evaluates the performance of the resulting approximation; it belongs to the evaluation stage.
- Gradient-based methods are a focused subset of a much larger design space, with linear gradient-descent methods receiving attention because they are simple, promising, and theoretically useful.
- V estimates the long-run reward from a state, while Q estimates the long-run reward from taking a particular action.
- Using V means comparing the values of states reached by actions; using Q means comparing the action-value estimates directly.
Key Takeaways
- Value prediction estimates total future reward rather than only immediate reward.
- Backups from reinforcement learning methods can supply training examples for function approximation.
- Function approximation supports generalization across many states and actions, while MSVE evaluates approximation performance.
- Gradient-based methods receive focused attention because they are promising and theoretically revealing, especially in their linear forms.
- State values evaluate states; action values evaluate particular actions, and the larger relevant estimate indicates the preferred choice.