Value Estimates in Reinforcement Learning
q∗(a) denotes the true value of action a.
Why Estimates Guide Decisions
An agent rarely knows the complete result of a decision immediately. Instead, it can treat each available choice as a forecast: which choice appears likely to produce the better long-term outcome? A value estimate provides that forecast by predicting the total reward the agent can accumulate in the future.
True Values and Learning Estimates
In a bandit problem, two closely related quantities describe an action. The true value of action a is written q∗(a). An estimate of that true value at time t is written Qt(a). The two expressions concern the same action, identified by a, but they are not the same quantity. q∗(a) is the value being estimated, whereas Qt(a) is the agent's estimate at a particular time.
What Time Adds to an Estimate
The subscript t in Qt(a) tells you when the estimate is being considered. It does not identify a different action; the action is still identified by a. It identifies the time of the estimate. As the agent gains experience, its estimate can be considered at a later time, so the notation distinguishes the estimate at one time from an estimate at another time.
What do you think happens?
An expression changes from Qt(a) to an estimate at a later time. What does the changed time notation tell you?
Reveal answer
Answer: The estimate is associated with a later time.
The subscript identifies the time of an estimate. The symbol a continues to identify the action, while q∗(a) remains the true value being estimated.
States, Actions, and Future Reward
A value estimate is a long-run prediction of total future reward. The distinction between state and action values is about what is being evaluated. V assigns an estimate to a state: it describes how promising it is for the agent to be in that situation. Q assigns an estimate to a specific action: it describes how promising it is to take that particular action. Neither description is limited to the next immediate result.
| Estimate | What it evaluates | What the prediction concerns |
|---|---|---|
| V | A state | Total future reward expected from being in that state |
| Q | A specific action | Total future reward expected from taking that action |
Two Ways to Prefer a Decision
Comparing Two Decisions
An agent is considering two actions. One action leads to state A, whose state value estimate is 8. The other action leads to state B, whose state value estimate is 3. Which action is preferred when the agent uses state value estimates?
Identify the relevant estimates: The first decision leads to state A with estimate 8. The second decision leads to state B with estimate 3.
Compare the predictions: The estimates represent predicted total future reward, not merely immediate rewards. Since 8 is larger than 3, state A has the larger predicted total future reward.
Select the preferred decision: Using state value estimates, the agent prefers the action that leads to state A.
The action leading to state A is preferred because its relevant state value estimate is 8, which is greater than 3.
The same preference can be formed through action values, but the comparison is made in a different place. With V, the agent compares the values of the states reached by the available actions. With Q, the agent compares the estimates already attached to the actions themselves. In either case, the preferred choice is associated with the larger predicted total future reward.
What do you think happens?
Suppose two available actions have action-value estimates Q(action 1) = 6 and Q(action 2) = 4. Which action is preferred when choosing through action values?
Reveal answer
Answer: Action 1
When using action values, the agent favors the action with the highest action-value estimate. Here, 6 is greater than 4.
Mistakes with Value Notation
Treating q∗(a) and Qt(a) as the same quantity.
q∗(a) denotes the true value, while Qt(a) denotes an estimate of that true value at time t.
Fix:
Use the star to recognize the true value and the subscript t to recognize the time-specific estimate.Interpreting t as part of the action name.
The symbol a identifies the action. The subscript t identifies when the estimate is being considered.
Fix:
Read Qt(a) as the estimate at time t for action a.Confusing a state value with an action value.
V assigns an estimate to a state, while Q assigns an estimate to a specific action.
Fix:
Ask what is being evaluated: the situation represented by a state, or a particular action.Treating a value estimate as an immediate reward.
The estimate is a prediction of total future reward available from the state, not a description limited to the next immediate result.
Fix:
Interpret the estimate as a long-run prediction.
Check Your Interpretation
For each statement, identify whether it describes q∗(a), Qt(a), V, or Q. Then explain what future reward the quantity is predicting: the reward associated with a state generally, or the reward associated with a particular action.
Hints
- Look for the star when identifying the true action value.
- Look for the subscript t when identifying an estimate at a particular time.
- V evaluates a state; Q evaluates a specific action.
- Value estimates concern total future reward, not only the next immediate result.
- Identify whether the notation refers to a true value or an estimate.
- If it is an estimate, inspect the time index.
- Determine whether the quantity evaluates a state or a specific action.
- Compare the relevant estimates when deciding which choice is preferred.
- Interpret the larger estimate as the larger predicted total future reward.
Essential Distinctions
- q∗(a) denotes the true value of action a.
- Qt(a) denotes the estimate at time t of q∗(a); a identifies the action and t identifies the time of the estimate.
- A value estimate is a long-run prediction of total future reward.
- V evaluates a state generally, while Q evaluates a particular action.
- State-value choice compares the values of states reached by actions; action-value choice compares estimates attached directly to actions.
Key Takeaways
- q∗(a) is the true value of action a, while Qt(a) is the estimate of that value at time t.
- The time index t identifies when the estimate is being considered, not a different action.
- Value estimates predict total future reward over the long run.
- V describes the value of a state, whereas Q describes the value of taking a particular action.
- An agent prefers the choice associated with the larger relevant value estimate.