On-policy Monte Carlo Control Methods
ε-greedy balances preference for the current best-known action with random action selection.
The Exploration Problem
A control method should favor actions that currently have the highest estimated action value. However, an on-policy method does not permanently choose only that action. The policy used to make decisions remains soft, so every available action keeps a positive chance of being selected. ε-greedy action selection combines these two requirements: it strongly favors the current best-known action while retaining a probability of trying an action at random.
The central idea is balance: estimated value determines which action is favored, while ε provides continued random-selection probability for every available action.
Splitting Probability Among Actions
An ε-greedy policy starts with the current estimated action values. The action, or actions, with the maximal estimated value are called greedy actions. Each available action receives a random-selection share of ε divided by the number of available actions. The greedy action also receives the remaining bulk of the probability. Therefore, a nongreedy action receives only its random-selection share, while the greedy action receives that same share plus the remaining preference for the best-known action.
A Four-Action Calculation
Four actions with ε equal to 0.2
Suppose four actions are available. One action is currently greedy, and ε is 0.2. Determine the selection probability for the greedy action and for each nongreedy action.
Divide the random share: The random-selection probability is divided equally among four available actions. Each action receives 0.2 divided by 4, which is 0.05.
Assign the nongreedy probabilities: Each nongreedy action receives its random-selection share, so each has probability 0.05.
Assign the greedy probability: The greedy action receives the remaining bulk, 0.8, plus its own random share, 0.05. Its total probability is 0.85.
Check the total: The four probabilities are 0.85, 0.05, 0.05, and 0.05. Together they account for the complete probability distribution.
The greedy action is selected with probability 0.85, and each of the three nongreedy actions is selected with probability 0.05.
| Action category | Number of actions | Probability per action |
|---|---|---|
| Greedy | 1 | 0.85 |
| Nongreedy | 3 | 0.05 each |
Generated calculation using four available actions and ε equal to 0.2.
ε-soft and ε-greedy Policies
| Policy class | What it requires | How specifically it distributes probability |
|---|---|---|
| ε-soft | Every available action has a positive chance of being selected. | It describes a requirement, not one unique distribution. |
| ε-greedy | Every available action remains selectable while the highest-estimated-value action is favored. | Each action receives its random-selection share; the greedy action receives the remaining bulk as well. |
ε-soft is the broader description. It says that every available action must retain a positive chance of selection, but it does not identify one unique probability distribution. ε-greedy is a particular way to meet that requirement: give every action the same random-selection share, then place the remaining bulk on the greedy action. In this sense, ε-greedy policies are the ε-soft policies that stay closest to greedy behavior while preserving exploration.
Improving the Acting Policy
On-policy control uses the current policy to make decisions and then improves that same policy from the action-value estimates it is building. The policy is not replaced by a permanently greedy rule. Instead, it remains soft: every action stays possible, while the action with the highest current estimated value receives the strongest preference. As the estimated action values change, the identity of the greedy action, and therefore the probability distribution, can change as well.
Tracing a Policy Change
Imagine that one action is currently identified as greedy because it has the highest estimated action value. The ε-greedy policy favors that action but still gives every available action a positive chance. After returns update the estimated action values, another action may become the one with the maximal estimate. The policy then changes its preference toward that newly greedy action, while continuing to keep every action selectable.
Common Probability Mistakes
Giving all of the remaining probability to the greedy action but forgetting its random-selection share.
The random share is assigned to every available action, including the greedy action.
Fix:
Add the greedy action's 0.05 random share to the remaining 0.8, giving it 0.85.Treating ε-soft and ε-greedy as exact synonyms.
ε-soft describes a requirement for positive selection probability, whereas ε-greedy specifies a particular distribution that satisfies that requirement.
Fix:
Describe ε-greedy as a member of the broader ε-soft policy class.Making the policy permanently greedy.
An on-policy method keeps the policy soft, so every available action retains a positive chance of selection.
Fix:
Favor the greedy action while preserving the random-selection share for every action.Choosing the greedy action without checking the current estimates.
The ε-greedy policy is built from the current estimated action values.
Fix:
Identify the action or actions with maximal estimated value before assigning the policy's preference.
Practice Check
Five actions are available, one action is currently greedy, and ε is 0.25. Determine the random-selection share for each action and the total selection probability for the greedy action.
Hints
- Divide ε by the number of available actions to find each action's random-selection share.
- Each nongreedy action receives only that share.
- The greedy action receives the remaining bulk plus its own random-selection share.
What do you think happens?
For five actions and ε equal to 0.25, what are the probabilities?
Reveal answer
Answer: The greedy action has probability 0.80, and each nongreedy action has probability 0.05.
The random-selection share is 0.25 divided by 5, which is 0.05 for every action. The greedy action also receives the remaining bulk of 0.75, giving it 0.80 in total.
Key Takeaways
- An ε-greedy policy favors the action with the highest current estimated action value while keeping every action selectable.
- Every available action receives a random-selection share of ε divided by the number of available actions.
- The greedy action receives its random-selection share plus the remaining bulk of probability.
- ε-soft is a broad policy requirement; ε-greedy is a specific probability structure within that class.
- On-policy control improves the same soft policy that selects actions, so updated action values can change which action receives the strongest preference.
Key Takeaways
- ε-greedy balances exploitation of the current best estimate with random action selection.
- Nongreedy actions receive the minimum random-selection share, while the greedy action receives that share plus the remaining bulk.
- ε-soft describes the positive-probability requirement; ε-greedy describes one particular way to distribute probability.
- On-policy control updates and improves the same policy that generated the actions.
- To calculate probabilities, first divide ε by the number of available actions, then add the remaining bulk to the greedy action.