Value Estimation in Reinforcement Learning
A reward signal describes immediate desirability, while a value function describes long-term desirability.
The tempting immediate choice
Imagine an agent choosing between two situations. One situation provides something desirable immediately. The other appears less attractive at first but tends to lead to better outcomes later. If the agent considers only the next moment, it chooses the first situation. Reinforcement learning needs a way to represent the longer-term consequence of beginning from each situation. That is the role of a value function.
Reward and value are different
A reward signal describes immediate desirability. It is supplied directly by the environment and represents what the agent receives on the current transition. A value estimate describes long-term desirability. It is a long-run prediction of the total future reward the agent can accumulate from a state or from taking a particular action.
Reading a state value
The value of a state estimates the total reward the agent can expect to accumulate in the future from that state. It answers a forecasting question: if the agent is in this situation, how much future reward does it expect to receive?
The estimate is not limited to the reward visible right now. It represents the future reward expected from the state. This is why an agent can prefer a situation that is initially less attractive: the situation may be more promising when its later consequences are considered.
Reading an action value
V assigns an estimate to a state, while Q assigns an estimate to a specific action. An action value describes the total future reward expected from taking that particular action. Both are long-run predictions rather than descriptions limited to the next immediate result.
The distinction is about what is being evaluated. V asks how promising it is for the agent to be in a situation. Q asks how promising it is to take a particular action. The first estimate is attached to a state; the second is attached to an action.
Comparing decision routes
Choosing through state values
An agent is considering two actions. One action leads to state A, whose value estimate is 8. The other leads to state B, whose value estimate is 3. Which action is preferred when the agent uses state value estimates?
Identify the compared quantities: The agent compares the values of the states reached by the available actions, not the immediate rewards of those actions.
Compare the estimates: State A has an estimate of 8, while state B has an estimate of 3.
Select the larger estimate: The action leading to state A is preferred because 8 is the larger prediction of total future reward.
The agent prefers the action that leads to state A. The numbers 8 and 3 are predictions of future total reward, not immediate rewards.
| Decision viewpoint | What is estimated | How the choice is guided |
|---|---|---|
| State values | The desirability of a state | Favor the action that leads to the highest-valued state |
| Action values | The desirability of a specific action | Favor the action with the highest action-value estimate |
Why estimation is difficult
Rewards come directly from the environment, so identifying an immediate reward signal is comparatively straightforward. Values are different because they must be inferred from sequences of observations gathered across the agent's lifetime. The agent must repeatedly estimate and re-estimate what a state is worth as it encounters more observations and learns about the outcomes that tend to follow.
Value estimation is central because an agent rarely knows the complete result of a decision immediately. Value estimates provide a forecast of which available choice appears to lead to the better long-term outcome.
Check your reasoning
An agent can choose between two situations. Situation A gives a larger reward immediately but tends to lead to less rewarding future states. Situation B gives a smaller reward immediately but tends to lead to more rewarding future states. Which situation may have the higher value, and why?
Hints
- Separate the current reward from the prediction about future rewards.
- Ask which situation tends to lead to the greater total future reward.
What do you think happens?
The action leading to state A has a state value estimate of 8. The action leading to state B has a state value estimate of 3. Which action is preferred when choosing through state values?
Reveal answer
Answer: The action leading to state A.
When using state values, the agent favors the action that leads to the highest-valued state. The estimate of 8 is larger than 3, and both numbers predict total future reward rather than only the next reward.
Mistakes to avoid
Treating a value estimate as the immediate reward.
The estimate is a prediction of total future reward, while the reward is supplied directly by the environment as an immediate signal.
Fix:
Describe the reward as immediate and the value as a long-run prediction.Assuming that a low immediate reward means the state has low value.
That state may tend to lead to future states offering greater rewards.
Fix:
Consider the total future reward predicted from the state.Confusing a state value with an action value.
V describes a state, whereas Q describes a specific action.
Fix:
Use V for the value of a state and Q for the value of taking a particular action.Comparing the wrong objects when selecting an action.
The state-value route compares the values of states reached by actions; the action-value route compares estimates attached directly to the actions.
Fix:
State clearly whether the decision is based on resulting-state values or available-action values.
Key takeaways
- A reward signal describes immediate desirability, while a value function describes long-term desirability.
- A state value estimates the total future reward expected from a state.
- An action value estimates the total future reward expected from taking a specific action.
- A state with a low immediate reward can still have high value when it tends to lead to more rewarding future states.
- Value estimates must be inferred from observation sequences and repeatedly updated as the agent learns about later outcomes.
Key Takeaways
- Immediate rewards describe what the environment supplies now; value estimates predict total future reward.
- State values evaluate situations, while action values evaluate particular choices.
- An agent can prefer a lower immediate reward when the associated state or action promises better future outcomes.
- Value estimation is challenging because it must be learned from sequences of observations and updated over time.