Optimal Value Functions
Optimality describes the highest-value policy, not the computational effort required to find it.
The Ideal and the Practical Agent
In reinforcement learning, it is natural to aim for the policy that achieves the highest value. That policy is called optimal. The important qualification is that optimality describes the strongest result, not the amount of computation required to discover it. An agent may have a clear description of its environment and still be unable to calculate an optimal action before the next decision is required.
Optimality is both a theoretical target and a standard for judging learning methods. It is not generally a policy that an agent can simply calculate and use.
Why Knowing the Environment Is Not Enough
An environment model provides information about how the environment behaves. Computation must then use that information to evaluate possible choices, compare their values, and select an action. The difficult part may be the amount of computation needed to generate the final policy, even when the environment dynamics are completely and accurately known.
A single time step creates a practical boundary. The agent observes a state, evaluates possibilities, and must select an action. Computation that is not completed before the decision is required cannot improve that immediate decision. Therefore, the agent is constrained not only by what it knows, but also by what it can compute quickly enough.
Chess illustrates the distinction. Even large, specialized computers may not compute optimal moves. This does not mean that chess has no useful policies or that strong play is impossible. It means that substantial computational resources do not guarantee that the mathematically optimal policy can be computed.
Reading an Optimal Value
An optimal value function evaluates the best achievable outcome from a state. In other words, v∗(s) answers this question: starting in state s, what is the maximum value that can be obtained if the available choices are followed optimally?
The value belongs to the state, not to one named first action. To interpret v∗(s), imagine starting at s and allowing the best available decisions to be made from there. The result is the highest value that can be achieved from that state.
Interpreting a State Value in Golf
What does an optimal value assigned to a golf state mean?
Identify the state: The state describes the current golfing situation, such as the ball's position and the strokes available.
Consider the available choices: The optimal value considers the available actions and selects choices that achieve the best possible result.
Interpret the value: The value is the best achievable outcome from that state. It is not a commitment to one particular named first stroke.
An optimal value function tells us how good the state is when the agent can choose the best available actions from that point onward.
The word optimal means that the value reflects the best result available after considering the choices from the state.
State Values and Action Values
An optimal value function evaluates a state before fixing a particular first action. An optimal action-value function evaluates a state after a particular first action has been chosen. The difference is the object being evaluated: v∗(s) concerns the state, while q∗(s, a) concerns the state-action pair.
The Driver Commitment
How should q∗(s, driver) be interpreted in the golf example?
Fix the first action: The first stroke is already committed to being a driver stroke.
Continue optimally: After that first stroke, later strokes can be selected optimally from the available choices of driver or putter.
Read the action value: q∗(s, driver) is the value of the state under that first-action commitment.
The optimal action-value function evaluates what can be achieved from a state when a particular first action has already been selected.
Bellman Optimality Relationships
Bellman optimality equations characterize optimal values by expressing the value of each state in terms of the choices available from that state. They create state-by-state relationships: the value assigned to a state must be consistent with the best choices available there and with the values of the states that can follow.
The recycling robot example makes the structure small enough to see clearly. Its state space contains two states, high and low, so the Bellman optimality description contains one equation for the high state and one equation for the low state. Each equation checks whether the value assigned to its state agrees with the choices available there.
It is useful to read these equations as mutual checks. The value of the high state must be consistent with the choices available in high, and the value of the low state must be consistent with the choices available in low. Solving the relationships gives the optimal values for the states.
From Values to an Optimal Policy
Once the optimal value function has been obtained, an optimal policy selects actions that achieve the maximum value from each state. The value function therefore supports policy construction: evaluate the available actions, identify those associated with the highest value, and choose one of them.
Reading a Gridworld Policy
What does an optimal policy do after the gridworld's optimal values have been solved?
Start with a state: The solved optimal value function assigns a value to the current gridworld state.
Compare actions: The possible actions from that state are compared according to the values they can achieve.
Keep maximum-valued actions: The policy identifies actions that achieve the maximum value from the state.
The corresponding optimal policy is obtained from the solved values. If several actions achieve the same maximum, all of them can be optimal.
An optimal policy need not specify only one action in every state. Multiple actions can be optimal when they achieve the same maximum value.
Mistakes in Interpreting Optimality
Assuming that a perfect environment model makes the optimal policy easy to compute.
Knowledge supplies information, but the final policy may still require extreme computation.
Fix:
Treat model accuracy and computational feasibility as separate questions.Treating optimality as a policy that every agent can calculate during ordinary operation.
Optimality identifies the highest-value policy; it does not guarantee that the policy can be produced in practice.
Fix:
Use optimality as a theoretical ideal and a benchmark for judging methods.Confusing the value of a state with the value of a state-action pair.
The action-value function includes a commitment to the particular first action.
Fix:
Use v∗(s) for the best value from a state and q∗(s, a) for the value after selecting a first action.Assuming an optimal policy must contain exactly one action in each state.
More than one action may achieve the maximum value.
Fix:
Allow every action tied for the maximum to be part of the optimal policy.
Practice: Separate the Questions
For each statement, decide whether it concerns an optimal value function, an optimal action-value function, computational feasibility, or policy construction: a value assigned to a state before a first action is fixed; a value assigned after the first action is driver; the time available before the next decision; and the actions that achieve the maximum value in a gridworld cell.
Hints
- Ask whether the statement evaluates a state or a state-action pair.
- Ask whether it describes the best theoretical result or the computation needed to find it.
- For policy construction, look for actions that achieve the maximum.
- Optimal value functions describe the best achievable outcome from states. Bellman optimality equations characterize those values through the available choices and resulting state values. An optimal policy selects actions that achieve the maximum value, with ties allowed. An optimal action-value function differs because it evaluates a state after a particular first action has been fixed. Finally, the theoretical optimum can remain difficult or impossible to compute within the time available for a real decision.
Key Takeaways
- An optimal policy is a policy that achieves the highest value, but finding it may require impractical computation.
- An optimal value function gives the best achievable value from a state without fixing one named first action.
- An optimal action-value function evaluates a state-action pair after a particular first action has been selected.
- Bellman optimality equations express consistency relationships among optimal values across states.
- An optimal policy can be obtained by selecting actions that achieve the maximum value, including multiple tied actions.