Concepts / Value Functions and Policies

Value Functions and Policies

Monte Carlo methods use sample episodes as experience.

  • Programming

From Experience to Evaluation

A policy tells an agent how to behave, but it does not by itself tell us how promising a situation is. Value functions provide that evaluation. They describe the expected return associated with a state or with a state-action pair when the agent follows a particular policy. Monte Carlo methods estimate these values from sample episodes, using what actually happens as experience.

The central distinction is between evaluating a particular policy and finding the best performance any policy could achieve.

ends withupdatesupdatesSample episodestates, actions, rewardsEventual outcomecompleted returnState estimatesvisited statesAction estimatesvisited state-action pairs
How does an episode's observed experience flow into updated estimates for the states or state-action pairs it visited?

Monte Carlo and Dynamic Programming

Monte Carlo and Dynamic Programming methods both address value functions and policies, but they can use different kinds of information. Monte Carlo methods learn from experience represented as sample episodes. They can therefore learn through interaction without an explicit model of environment dynamics. Dynamic Programming relies on a fuller description of the environment, including transition and reward dynamics, and uses estimated values in its updates.

learns frominformsrequiressupportsMonte Carlosample episodesObserved experiencecompleted episodeValue estimatefrom sampled returnDynamic Programmingenvironment modelTransition and rewarddynamicsknown descriptionValue estimateuses successor estimates
How do Monte Carlo methods learn from sampled episodes while Dynamic Programming uses a model of transition and reward dynamics and bootstraps from value estimates?

Generated example: Imagine a learning situation with states A, B, and C. An episode begins at A, passes through B, and ends at C. A Monte Carlo method can use this completed episode as experience. It does not need an explicit table describing every possible transition in the environment in order to use the observed sequence and its eventual outcome.

Selective Attention and Markov Effects

Monte Carlo methods can concentrate computation on a selected subset of states. If an application contains a large collection of states but only a small region matters for the current investigation, experience can be directed toward that region. This is selective attention: the method avoids spending the same effort accurately evaluating states outside the immediate area of interest. It does not mean that all other states have been evaluated equally well.

Monte Carlo methods also avoid bootstrapping. Their update for a state does not use the current value estimate of a successor state. Instead, the method uses the completed sampled return from an episode. Because the update does not depend on a successor-state value estimate, violations of the Markov property may harm Monte Carlo methods less.

updates fromupdates fromCurrent statebeforeSuccessor estimateused by updateCurrent stateafterCompleted returnfrom sampled episode
What changes when a method estimates a value from a complete sampled return rather than from another potentially inaccurate estimate?

When comparing methods, ask what information the update depends on. A method that depends on a successor-state estimate can inherit problems from that estimate. A Monte Carlo update based on a completed sampled return avoids that particular dependency.

State and Action Evaluation

A policy's state-value function measures the expected return associated with a state when the agent follows that policy. A policy's action-value function measures the expected return associated with a state-action pair when the agent follows that policy.

evaluatesevaluatesStateexpected return underpolicyState valuestate evaluationState-action pairaction choice includedAction valuepair evaluation
What does a policy's state-value function measure for a state, and what does its action-value function measure for each state-action pair?

Separating a Situation from an Action

An agent is in state B and has to decide what to do next. What is the difference between evaluating B and evaluating a particular action taken in B?

Evaluate the state: The state-value function asks how promising state B is when the agent behaves according to the selected policy.

Evaluate a state-action pair: The action-value function asks how promising the particular pair consisting of state B and one selected action is when the policy is followed afterward.

Use the distinction: The state evaluation summarizes the situation, while the action evaluation includes the immediate choice that is being considered.

A state value evaluates a state under a policy; an action value evaluates a state-action pair under that policy.

Optimal Values and Policy Choice

Ordinary value functions evaluate expected return under a particular policy. Optimal value functions ask a different question: what is the largest expected return available across policies? The optimal value functions therefore represent the best performance that can be achieved from states or state-action pairs.

Optimal values can be used to construct an optimal policy through greedy choices. The policy selects actions that maximize the optimal action-value function. Several policies can be optimal because different policies may achieve the same optimal value, even though the optimal value functions themselves are unique.

supportssupportsachievesachievesOptimal actionvaluesbest available returnsOptimal policy 1greedy choiceOptimal valuesame best performanceOptimal policy 2also greedy where tied
How does the optimal action-value function identify an optimal policy, and how can several policies achieve the same optimal value?

Bellman Optimality as a Constraint

Bellman optimality equations provide the consistency conditions for optimal values. They constrain the values so that they agree with optimal behavior: the best available immediate action and the corresponding continuation value determine the value assigned to a situation. Once the optimal values have been found, policy selection becomes relatively easy because a policy can choose actions greedily with respect to those values.

considersselectsconstrainssupportsState orstate-actionsituationAvailable actionscompare possibilitiesBest actionlargest available returnOptimal valueconsistent with optimalbehaviorGreedy policyselects best action
How does the Bellman optimality backup choose the best immediate action and continuation value to define or improve an optimal value function?

The equations do more than produce numbers. They make the values mutually consistent with optimal behavior; those values then make greedy policy selection possible.

Mistakes in Reasoning

  • Treating a policy as an evaluation

    A policy specifies behavior, while a value function evaluates expected return under that behavior.

    Fix: Keep the two roles separate: policy for choosing behavior, value function for measuring its expected return.

  • Assuming Monte Carlo methods require a complete environment model

    Monte Carlo methods can learn from sample episodes and interaction without an explicit model of environment dynamics.

    Fix: Use observed experience, simulation, or a sample model when a complete model is difficult to construct.

  • Confusing a state value with an action value

    A state-value function evaluates a state; an action-value function evaluates a state-action pair.

    Fix: Check whether the action choice is part of the object being evaluated.

  • Assuming optimal values imply one unique optimal policy

    Several policies may achieve the same optimal value.

    Fix: Allow for multiple policies that are optimal even when the optimal values are unique.

  • Calling a successor-value update a complete-return update

    Monte Carlo methods avoid bootstrapping because their updates do not use successor-state value estimates.

    Fix: Identify whether the update depends on a completed sampled return or on another estimated value.

Check Your Understanding

MEDIUM

An investigation concerns only a small region of a very large state set. Explain why a Monte Carlo approach may be attractive. Then distinguish the state-value question from the action-value question for one state in that region. Finally, explain how an optimal value function could guide action selection and why more than one optimal policy might still exist.

Hints
  • Begin with the kind of information Monte Carlo methods use.
  • For the two value questions, ask whether a particular action is included.
  • For multiple policies, focus on whether different choices can achieve the same optimal value.

Key Takeaways

  1. Monte Carlo methods learn from sample episodes and can interact with an environment without an explicit model of its dynamics.
  2. Simulation and sample models are useful when constructing a complete transition-probability model is difficult.
  3. Monte Carlo computation can focus on a selected subset of states, and avoiding bootstrapping can make violations of the Markov property less harmful.
  4. A policy's state-value function evaluates states, while its action-value function evaluates state-action pairs.
  5. Optimal value functions represent the largest expected return across policies; Bellman optimality equations constrain those values, and greedy choices with respect to them identify optimal policies.
  6. Multiple policies may be optimal even though the optimal value functions are unique.

Key Takeaways

  • Monte Carlo methods use sampled episodes as experience instead of requiring a complete explicit environment model.
  • They can use simulation or sample models, focus computation on selected states, and avoid successor-value bootstrapping.
  • Value functions measure expected return under a particular policy for states or state-action pairs.
  • Optimal value functions describe the best performance available across policies and support greedy policy construction.
  • Bellman optimality equations enforce consistency with optimal behavior, while several different policies may still achieve the same optimal value.