Concepts / On-policy Monte Carlo Control Methods

On-policy Monte Carlo Control Methods

ε-greedy balances preference for the current best-known action with random action selection.

  • Programming

The Exploration Problem

A control method should favor actions that currently have the highest estimated action value. However, an on-policy method does not permanently choose only that action. The policy used to make decisions remains soft, so every available action keeps a positive chance of being selected. ε-greedy action selection combines these two requirements: it strongly favors the current best-known action while retaining a probability of trying an action at random.

The central idea is balance: estimated value determines which action is favored, while ε provides continued random-selection probability for every available action.

Splitting Probability Among Actions

An ε-greedy policy starts with the current estimated action values. The action, or actions, with the maximal estimated value are called greedy actions. Each available action receives a random-selection share of ε divided by the number of available actions. The greedy action also receives the remaining bulk of the probability. Therefore, a nongreedy action receives only its random-selection share, while the greedy action receives that same share plus the remaining preference for the best-known action.

Greedy actionrandom share plus remainingbulkNongreedy action 1ε divided by action countNongreedy action 2ε divided by action count
How is the total probability split between the current greedy action and each nongreedy action?

A Four-Action Calculation

Four actions with ε equal to 0.2

Suppose four actions are available. One action is currently greedy, and ε is 0.2. Determine the selection probability for the greedy action and for each nongreedy action.

Divide the random share: The random-selection probability is divided equally among four available actions. Each action receives 0.2 divided by 4, which is 0.05.

Assign the nongreedy probabilities: Each nongreedy action receives its random-selection share, so each has probability 0.05.

Assign the greedy probability: The greedy action receives the remaining bulk, 0.8, plus its own random share, 0.05. Its total probability is 0.85.

Check the total: The four probabilities are 0.85, 0.05, 0.05, and 0.05. Together they account for the complete probability distribution.

The greedy action is selected with probability 0.85, and each of the three nongreedy actions is selected with probability 0.05.

Action categoryNumber of actionsProbability per action
Greedy10.85
Nongreedy30.05 each

Generated calculation using four available actions and ε equal to 0.2.

Greedy action0.85Nongreedy action 10.05Nongreedy action 20.05Nongreedy action 30.05
Given ε equal to 0.2 and four available actions, what probability does each greedy and nongreedy action receive?

ε-soft and ε-greedy Policies

Policy classWhat it requiresHow specifically it distributes probability
ε-softEvery available action has a positive chance of being selected.It describes a requirement, not one unique distribution.
ε-greedyEvery available action remains selectable while the highest-estimated-value action is favored.Each action receives its random-selection share; the greedy action receives the remaining bulk as well.

ε-soft is the broader description. It says that every available action must retain a positive chance of selection, but it does not identify one unique probability distribution. ε-greedy is a particular way to meet that requirement: give every action the same random-selection share, then place the remaining bulk on the greedy action. In this sense, ε-greedy policies are the ε-soft policies that stay closest to greedy behavior while preserving exploration.

specific member ofε-softpositive chance for everyactionε-greedyrandom shares plus greedybulk
What constraint does every ε-soft policy satisfy, and what additional probability structure makes a policy specifically ε-greedy?

Improving the Acting Policy

On-policy control uses the current policy to make decisions and then improves that same policy from the action-value estimates it is building. The policy is not replaced by a permanently greedy rule. Instead, it remains soft: every action stays possible, while the action with the highest current estimated value receives the strongest preference. As the estimated action values change, the identity of the greedy action, and therefore the probability distribution, can change as well.

selects actionsproduces returnschanges preferenceacts againCurrent policysoft action selectionEnvironmentreturnsAction valuescurrent estimatesImproved policynew preference
How do actions selected by the current policy produce returns that update action values and improve that same policy?

Tracing a Policy Change

Imagine that one action is currently identified as greedy because it has the highest estimated action value. The ε-greedy policy favors that action but still gives every available action a positive chance. After returns update the estimated action values, another action may become the one with the maximal estimate. The policy then changes its preference toward that newly greedy action, while continuing to keep every action selectable.

updated estimatesupdated estimatesAction AgreedyAction AnongreedyAction BnongreedyAction Bgreedy
What changes when updated action-value estimates identify a different action as best?

Common Probability Mistakes

  • Giving all of the remaining probability to the greedy action but forgetting its random-selection share.

    The random share is assigned to every available action, including the greedy action.

    Fix: Add the greedy action's 0.05 random share to the remaining 0.8, giving it 0.85.

  • Treating ε-soft and ε-greedy as exact synonyms.

    ε-soft describes a requirement for positive selection probability, whereas ε-greedy specifies a particular distribution that satisfies that requirement.

    Fix: Describe ε-greedy as a member of the broader ε-soft policy class.

  • Making the policy permanently greedy.

    An on-policy method keeps the policy soft, so every available action retains a positive chance of selection.

    Fix: Favor the greedy action while preserving the random-selection share for every action.

  • Choosing the greedy action without checking the current estimates.

    The ε-greedy policy is built from the current estimated action values.

    Fix: Identify the action or actions with maximal estimated value before assigning the policy's preference.

Practice Check

EASY

Five actions are available, one action is currently greedy, and ε is 0.25. Determine the random-selection share for each action and the total selection probability for the greedy action.

Hints
  • Divide ε by the number of available actions to find each action's random-selection share.
  • Each nongreedy action receives only that share.
  • The greedy action receives the remaining bulk plus its own random-selection share.

What do you think happens?

For five actions and ε equal to 0.25, what are the probabilities?

  • Greedy: 0.75; each nongreedy: 0.05
  • Greedy: 0.80; each nongreedy: 0.05
  • Greedy: 0.25; each nongreedy: 0.15
Reveal answer

Answer: The greedy action has probability 0.80, and each nongreedy action has probability 0.05.

The random-selection share is 0.25 divided by 5, which is 0.05 for every action. The greedy action also receives the remaining bulk of 0.75, giving it 0.80 in total.

Key Takeaways

  1. An ε-greedy policy favors the action with the highest current estimated action value while keeping every action selectable.
  2. Every available action receives a random-selection share of ε divided by the number of available actions.
  3. The greedy action receives its random-selection share plus the remaining bulk of probability.
  4. ε-soft is a broad policy requirement; ε-greedy is a specific probability structure within that class.
  5. On-policy control improves the same soft policy that selects actions, so updated action values can change which action receives the strongest preference.

Key Takeaways

  • ε-greedy balances exploitation of the current best estimate with random action selection.
  • Nongreedy actions receive the minimum random-selection share, while the greedy action receives that share plus the remaining bulk.
  • ε-soft describes the positive-probability requirement; ε-greedy describes one particular way to distribute probability.
  • On-policy control updates and improves the same policy that generated the actions.
  • To calculate probabilities, first divide ε by the number of available actions, then add the remaining bulk to the greedy action.