Decision-Making with Value Functions
A value estimate is a long-run prediction of total future reward.
Forecasting Long-Term Reward
An agent rarely knows the complete result of a decision immediately. Instead, it can treat each available choice as a forecast: which choice appears to lead to the better long-term outcome? Value estimates provide this forecast. They predict the total reward the agent can accumulate in the future, rather than describing only the next immediate result.
A value estimate is a long-run prediction of total future reward.
A larger value estimate represents a more promising long-term outcome for the agent when the estimates are being compared for the same decision.
Choosing Through State Values
A state value estimate, written as V, assigns an estimate to a state. It describes how promising it is for the agent to be in that situation by predicting the total future reward expected from the state. To choose through state values, the agent considers where each available action leads, compares the values of the resulting states, and favors the action leading to the state with the highest value.
Comparing the States Reached
An agent is considering two actions. One action leads to State A, whose value estimate is 8. The other leads to State B, whose value estimate is 3. Which action is preferred when the agent uses state values?
Identify the states: The two actions lead to State A and State B.
Read the estimates: State A has an estimated value of 8, while State B has an estimated value of 3.
Compare the estimates: The estimate for State A is larger than the estimate for State B.
Select the action: The agent prefers the action that leads to State A because 8 is the larger predicted total future reward.
The action leading to State A is preferred.
State Estimates and Action Estimates
The difference between V and Q is what is being evaluated. V describes a state: how promising is it for the agent to be in this situation? Q describes a specific action: how promising is it to take this particular action? Both are long-run predictions of total future reward, but they attach the prediction to different parts of the decision.
| Estimate | What it evaluates | How it guides a choice |
|---|---|---|
| V | A state | Favor the action that leads to the highest-valued state |
| Q | A specific action | Favor the action with the highest action-value estimate |
Choosing Through Action Values
An action value estimate, written as Q, assigns an estimate to a specific action. It describes the total future reward expected from taking that particular action. When an agent uses action values, it does not first compare values attached to the states reached by the actions. Instead, it compares the Q estimates already attached to the available actions and favors the action with the highest estimate.
The state-value route and the action-value route differ only in where the comparison is made. With V, the agent compares the values of states reached by actions. With Q, the agent compares estimates already attached to the actions themselves. In either case, the estimate representing the larger predicted total future reward is the preferred one.
Reading a Value-Based Decision
What do you think happens?
An agent can choose between an action leading to State A with V equal to 8 and an action leading to State B with V equal to 3. If the agent uses state values, which action should it prefer?
Reveal answer
Answer: The action leading to State A
State values are compared for the states reached by the available actions. Because 8 is greater than 3, the action leading to State A has the more promising long-term forecast.
When reading a value-based decision, first identify whether the estimates belong to states or to actions. Then compare estimates of the same kind and select the choice associated with the largest relevant estimate.
Mistakes in Value Comparisons
Treating a value estimate as an immediate reward
The source example describes 8 as a prediction of the total reward available from the state over the future.
Fix:
Interpret the value as a long-run forecast, not as a reward limited to the next step.Saying that V and Q evaluate the same thing
V assigns an estimate to a state, whereas Q assigns an estimate to a specific action.
Fix:
Ask whether the estimate describes being in a state or taking a particular action.Comparing the wrong objects
The state-value route compares states reached by actions; the action-value route compares estimates attached directly to actions.
Fix:
Identify the estimate type first, then compare the matching states or actions.
Decision Practice
An agent has two available actions. Action North leads to a state with a value estimate of 6. Action South leads to a state with a value estimate of 2. Assume the agent is choosing through state values. Which action is preferred, and what does the larger number predict?
Hints
- Compare the value estimates of the two states reached by the actions.
- The preferred action is the one associated with the larger state-value estimate.
- Explain the larger number as a prediction of total future reward, not as an immediate reward.
- Value estimates help an agent make decisions by forecasting total future reward. V describes a state and guides the agent toward actions that lead to higher-valued states. Q describes a specific action and guides the agent toward actions with higher action-value estimates. The central question is always what is being evaluated: a state or an action. Once that is clear, the agent prefers the choice associated with the largest relevant long-term estimate.
Key Takeaways
- A value estimate is a long-run prediction of total future reward.
- V assigns an estimate to a state, while Q assigns an estimate to a specific action.
- Using state values means favoring actions that lead to the highest-valued states.
- Using action values means favoring actions with the highest action-value estimates.
- The preferred decision is associated with the largest relevant predicted total future reward.