Value Functions and Policies
Monte Carlo methods use sample episodes as experience.
From Experience to Evaluation
A policy tells an agent how to behave, but it does not by itself tell us how promising a situation is. Value functions provide that evaluation. They describe the expected return associated with a state or with a state-action pair when the agent follows a particular policy. Monte Carlo methods estimate these values from sample episodes, using what actually happens as experience.
The central distinction is between evaluating a particular policy and finding the best performance any policy could achieve.
Monte Carlo and Dynamic Programming
Monte Carlo and Dynamic Programming methods both address value functions and policies, but they can use different kinds of information. Monte Carlo methods learn from experience represented as sample episodes. They can therefore learn through interaction without an explicit model of environment dynamics. Dynamic Programming relies on a fuller description of the environment, including transition and reward dynamics, and uses estimated values in its updates.
Generated example: Imagine a learning situation with states A, B, and C. An episode begins at A, passes through B, and ends at C. A Monte Carlo method can use this completed episode as experience. It does not need an explicit table describing every possible transition in the environment in order to use the observed sequence and its eventual outcome.
Selective Attention and Markov Effects
Monte Carlo methods can concentrate computation on a selected subset of states. If an application contains a large collection of states but only a small region matters for the current investigation, experience can be directed toward that region. This is selective attention: the method avoids spending the same effort accurately evaluating states outside the immediate area of interest. It does not mean that all other states have been evaluated equally well.
Monte Carlo methods also avoid bootstrapping. Their update for a state does not use the current value estimate of a successor state. Instead, the method uses the completed sampled return from an episode. Because the update does not depend on a successor-state value estimate, violations of the Markov property may harm Monte Carlo methods less.
When comparing methods, ask what information the update depends on. A method that depends on a successor-state estimate can inherit problems from that estimate. A Monte Carlo update based on a completed sampled return avoids that particular dependency.
State and Action Evaluation
A policy's state-value function measures the expected return associated with a state when the agent follows that policy. A policy's action-value function measures the expected return associated with a state-action pair when the agent follows that policy.
Separating a Situation from an Action
An agent is in state B and has to decide what to do next. What is the difference between evaluating B and evaluating a particular action taken in B?
Evaluate the state: The state-value function asks how promising state B is when the agent behaves according to the selected policy.
Evaluate a state-action pair: The action-value function asks how promising the particular pair consisting of state B and one selected action is when the policy is followed afterward.
Use the distinction: The state evaluation summarizes the situation, while the action evaluation includes the immediate choice that is being considered.
A state value evaluates a state under a policy; an action value evaluates a state-action pair under that policy.
Optimal Values and Policy Choice
Ordinary value functions evaluate expected return under a particular policy. Optimal value functions ask a different question: what is the largest expected return available across policies? The optimal value functions therefore represent the best performance that can be achieved from states or state-action pairs.
Optimal values can be used to construct an optimal policy through greedy choices. The policy selects actions that maximize the optimal action-value function. Several policies can be optimal because different policies may achieve the same optimal value, even though the optimal value functions themselves are unique.
Bellman Optimality as a Constraint
Bellman optimality equations provide the consistency conditions for optimal values. They constrain the values so that they agree with optimal behavior: the best available immediate action and the corresponding continuation value determine the value assigned to a situation. Once the optimal values have been found, policy selection becomes relatively easy because a policy can choose actions greedily with respect to those values.
The equations do more than produce numbers. They make the values mutually consistent with optimal behavior; those values then make greedy policy selection possible.
Mistakes in Reasoning
Treating a policy as an evaluation
A policy specifies behavior, while a value function evaluates expected return under that behavior.
Fix:
Keep the two roles separate: policy for choosing behavior, value function for measuring its expected return.Assuming Monte Carlo methods require a complete environment model
Monte Carlo methods can learn from sample episodes and interaction without an explicit model of environment dynamics.
Fix:
Use observed experience, simulation, or a sample model when a complete model is difficult to construct.Confusing a state value with an action value
A state-value function evaluates a state; an action-value function evaluates a state-action pair.
Fix:
Check whether the action choice is part of the object being evaluated.Assuming optimal values imply one unique optimal policy
Several policies may achieve the same optimal value.
Fix:
Allow for multiple policies that are optimal even when the optimal values are unique.Calling a successor-value update a complete-return update
Monte Carlo methods avoid bootstrapping because their updates do not use successor-state value estimates.
Fix:
Identify whether the update depends on a completed sampled return or on another estimated value.
Check Your Understanding
An investigation concerns only a small region of a very large state set. Explain why a Monte Carlo approach may be attractive. Then distinguish the state-value question from the action-value question for one state in that region. Finally, explain how an optimal value function could guide action selection and why more than one optimal policy might still exist.
Hints
- Begin with the kind of information Monte Carlo methods use.
- For the two value questions, ask whether a particular action is included.
- For multiple policies, focus on whether different choices can achieve the same optimal value.
Key Takeaways
- Monte Carlo methods learn from sample episodes and can interact with an environment without an explicit model of its dynamics.
- Simulation and sample models are useful when constructing a complete transition-probability model is difficult.
- Monte Carlo computation can focus on a selected subset of states, and avoiding bootstrapping can make violations of the Markov property less harmful.
- A policy's state-value function evaluates states, while its action-value function evaluates state-action pairs.
- Optimal value functions represent the largest expected return across policies; Bellman optimality equations constrain those values, and greedy choices with respect to them identify optimal policies.
- Multiple policies may be optimal even though the optimal value functions are unique.
Key Takeaways
- Monte Carlo methods use sampled episodes as experience instead of requiring a complete explicit environment model.
- They can use simulation or sample models, focus computation on selected states, and avoid successor-value bootstrapping.
- Value functions measure expected return under a particular policy for states or state-action pairs.
- Optimal value functions describe the best performance available across policies and support greedy policy construction.
- Bellman optimality equations enforce consistency with optimal behavior, while several different policies may still achieve the same optimal value.