Concepts / Monte Carlo Policy Evaluation

Monte Carlo Policy Evaluation

Continual exploration is necessary for Monte Carlo policy evaluation of action values.

  • Programming

Why Coverage Matters

Monte Carlo policy evaluation estimates action-values from experience in complete episodes. That estimate can improve only when the state-action pairs being evaluated are encountered. If some pairs are never encountered, the method cannot continually learn from those pairs. The central requirement is therefore not merely to complete episodes, but to keep every relevant state-action pair reachable through experience.

Exploring Starts Across Episodes

The exploring-starts assumption gives every state-action pair a nonzero probability of being the starting pair of an episode. The method does not require every pair to be selected equally often. It requires only that no pair has zero probability of starting an episode.

can startcan startcan startState-action pair Anonzero start probabilityState-action pair Bnonzero start probabilityState-action pair Cnonzero start probabilityInfinitely manyepisodeseach pair is encounteredinfinitely often
How do nonzero starting probabilities provide continuing coverage of every state-action pair across infinitely many episodes?

Applying the exploring-starts condition

Suppose an environment has several state-action pairs, and each pair has a nonzero probability of being selected as the starting pair.

Check the probabilities: Verify that no state-action pair has zero probability of starting an episode.

Consider infinitely many episodes: Under the exploring-starts assumption, the nonzero probability for each pair guarantees that every pair is encountered infinitely often in the limit.

Connect coverage to evaluation: Because each pair continues to appear in experience, Monte Carlo evaluation can continue learning about the action-values of those pairs.

Exploring starts supplies the coverage needed for action-value evaluation in the limit of infinitely many episodes.

Two Sources of Exploration

Exploration sourceHow actions are encounteredWhere the coverage comes from
Exploring startsEpisodes begin with state-action pairs selected under a starting ruleThe starting rule gives every state-action pair a nonzero probability
Stochastic policyThe agent sometimes selects each action while interacting in a stateThe policy gives every action in each state a nonzero selection probability
suppliessuppliesExploring startsnonzero probability forevery starting pairStochastic policynonzero probability forevery actionState-action coveragethrough episode beginningsAction coveragethrough policy choices
What is the difference between coverage supplied by episode starting conditions and coverage supplied by action selection during interaction?

Exploring starts can be useful in an idealized setting, but it cannot generally be assumed during direct interaction with an environment. Real starting conditions may not be arranged to cover state-action pairs. In that situation, exploration must be handled through another mechanism, such as a stochastic policy that gives every action in each state a nonzero probability of selection.

From Episodes to Action-Values

Policy evaluation gathers many complete episodes while following the current policy. For each visited state-action pair, the experienced episode supplies an observed return. Those returns are used to estimate the action-value function for the current policy. One episode does not produce the exact action-value function; the approximate function approaches the true function asymptotically as many episodes are experienced.

containslead toinformExperienced episodecomplete interactionVisited state-actionpairschoices observed in theepisodeObserved returnsreturn associated with eachvisitAction-valueestimatesupdated from many episodes
How does an experienced episode move from visited state-action pairs to observed returns and then to updated action-value estimates?

Evaluating an initial policy

Consider an initial policy named π0. The agent follows π0 through many complete episodes.

Follow π0: The episodes are generated while the current policy π0 determines the agent's behavior.

Record experience: For each state-action choice encountered, the episode provides an associated return.

Estimate action-values: The returns from many episodes are used to estimate how valuable each state-action choice is when continuing to follow π0.

Evaluation produces an approximate action-value function for π0; with infinitely many suitable episodes, the action-value function is computed exactly under the stated idealized assumptions.

Evaluation and Improvement

followproducessupplies basis forconstructsevaluate againprogresses towardInitial policy π0arbitrary policyPolicy evaluationestimate action-values fromepisodesAction-value functionvalues for the currentpolicyPolicy improvementchoose maximalaction-valuesImproved policynext policyOptimal policyunder the statedassumptions
How does the process alternate between evaluating the current policy and improving it until the policy approaches optimality?

Policy evaluation changes the available knowledge about the current policy: experienced episodes are used to estimate its action-values. Policy improvement changes the policy itself: the action-values are compared, and a new policy is made greedy with respect to them. Improvement is not another round of episode averaging. It is a decision rule applied to the action-value function produced by evaluation.

generatesestimatedeterminesCurrent policyinput to evaluationEpisode experiencereturns from following thepolicyAction-value functionoutput of evaluationGreedy policyoutput of improvement
What changes during evaluation, what changes during improvement, and how are their inputs and outputs connected?

Common Misunderstandings

  • Assuming that completing episodes is enough for action-value evaluation.

    Monte Carlo evaluation needs experience for the pairs whose action-values are being estimated.

    Fix: Ensure continual exploration through exploring starts or, in direct interaction, through a stochastic policy with nonzero probability for every action in each state.

  • Interpreting exploring starts as equal-frequency starts.

    The requirement is only that every pair has a nonzero probability of starting an episode.

    Fix: Focus on nonzero probability, not equal probability.

  • Treating a brief exploratory phase as continual exploration.

    Exploration must remain available throughout the learning process for the stated action-value evaluation requirement.

    Fix: Use a mechanism that continues to provide access to the pairs being evaluated.

  • Mixing up evaluation and improvement.

    Evaluation estimates action-values from episodes, whereas improvement applies a decision rule to those estimates.

    Fix: First evaluate the current policy, then construct a policy that chooses maximal action-values.

  • Assuming that one episode gives the exact action-value function.

    Monte Carlo evaluation uses many episodes, and the approximate action-value function approaches the true function asymptotically.

    Fix: Treat individual returns as experience contributing to estimates gathered across episodes.

Practice the Cycle

MEDIUM

A current policy is followed through many complete episodes. The episodes provide returns for the state-action pairs encountered, and the resulting estimates show that one action has the maximal action-value at a particular state. Identify which operation has just occurred, which operation comes next, and what the next policy should do at that state.

Hints
  • Episode returns are used during policy evaluation.
  • The operation that uses maximal action-values is policy improvement.
  • The next policy should choose an action with the maximal estimated action-value.

What do you think happens?

If every state-action pair has a nonzero probability of starting an episode, what does the exploring-starts assumption guarantee in the limit of infinitely many episodes?

  • Every pair is selected equally often
  • Every pair is encountered infinitely often
  • Only the currently best pair is encountered
  • No pair needs to be encountered again
Reveal answer

Answer: Every pair is encountered infinitely often.

The guarantee depends on every pair having nonzero starting probability. Equal frequency is not required.

Learning Progression

  1. Begin with an arbitrary policy.
  2. Maintain coverage of state-action pairs through exploring starts or suitable stochastic policy behavior.
  3. Follow the current policy through many complete episodes.
  4. Use the experienced returns to estimate the current policy's action-value function.
  5. Improve the policy by choosing an action with maximal action-value at each state.
  6. Evaluate the new policy and repeat the evaluation-improvement cycle.
  7. Under the stated assumptions, the cycle ends with the optimal policy and optimal action-value function.
  1. Monte Carlo action-value evaluation requires continual exploration because every pair being evaluated must continue to receive experience. Exploring starts provide coverage by assigning every state-action pair a nonzero probability of beginning an episode, which guarantees infinite encounters in the limit of infinitely many episodes. In direct interaction, a stochastic policy can provide an alternative by assigning every action in each state a nonzero selection probability. Monte Carlo policy iteration separates evaluation from improvement: episodes estimate the current policy's action-values, and improvement makes the next policy greedy with respect to those values. Repeating these steps moves an arbitrary initial policy toward the optimal policy under the stated assumptions.

Key Takeaways

  • Continual exploration is necessary because action-values can be learned only from experience involving the state-action pairs being evaluated.
  • Exploring starts guarantee coverage by giving every state-action pair a nonzero probability of beginning an episode; over infinitely many episodes, every pair is encountered infinitely often.
  • A stochastic policy can provide exploration during direct interaction by giving every action in each state a nonzero probability of selection.
  • Policy evaluation estimates action-values from experienced episodes, while policy improvement constructs a greedy policy from those estimates.
  • Repeating evaluation and improvement begins with an arbitrary policy and, under the stated assumptions, ends with the optimal policy and optimal action-value function.