Monte Carlo Control without Exploring Starts
Exploring starts assumes that all actions can be explored at the beginning, but this assumption is considered unrealistic.
Why the Starting Assumption Fails
Monte Carlo control needs experience containing different actions. Exploring starts handles this by assuming that the process can begin with all actions being explored. That gives the method a convenient starting condition, but the condition may not be available in the setting being modeled. In a real environment, an agent may not be able to choose every possible initial state and first action.
From Initial Exploration to Repeated Choice
Removing exploring starts does not remove the need for experience with different actions. Instead, the general alternative is to continue selecting actions so that all actions are selected infinitely often. The emphasis moves from the beginning of an episode to the repeated decision process across episodes. Rather than requiring every action to appear immediately at the start, the method keeps giving every action opportunities to be selected as experience is collected.
Replacing a Convenient Start
An agent needs experience containing different actions, but the environment cannot be assumed to begin with every action explored.
Identify the limitation: The agent cannot rely on the initial condition to explore every action.
Move exploration into repeated decisions: The agent continues selecting actions across episodes so that all actions are selected infinitely often.
Preserve varied experience: Repeated selection gives the agent experience that includes different actions without requiring exploring starts.
The method replaces an unrealistic initial-condition assumption with continued action selection over repeated episodes.
The Policy Relationship
An on-policy method evaluates or improves the same policy that is used to make decisions. An off-policy method evaluates or improves a different policy from the one used to generate the data.
| Method | Policy used to make decisions | Policy evaluated or improved |
|---|---|---|
| On-policy | The policy at the center of decision-making | That same policy |
| Off-policy | The policy that generates the data | A different policy |
The On-Policy Route
Monte Carlo control without exploring starts uses the on-policy route. The policy that makes decisions remains at the center of the method. The agent does not assume that the initial condition will explore every action. Instead, its decision process continues selecting actions so that all actions are selected infinitely often. This supplies the different-action experience needed by Monte Carlo control while keeping the policy used for decisions as the policy being evaluated or improved.
Following One Method Through Time
What do you think happens?
An agent cannot guarantee that every action is explored at the beginning of an episode. What general change allows an on-policy Monte Carlo control method to proceed?
Reveal answer
Answer: Continue selecting actions so that all actions are selected infinitely often
This is the general alternative to the unrealistic exploring-starts assumption. The on-policy method keeps the decision-making policy at the center while continuing action selection across episodes.
A Conceptual Episode Sequence
Trace the design of an on-policy Monte Carlo control method when the initial condition cannot be assumed to explore every action.
Initial condition: The episode begins without assuming that all actions have been explored at the beginning.
Decision making: The agent selects actions using the policy that is central to the method.
Repeated selection: Across repeated episodes, the action-selection design continues giving every action opportunities to be selected, so all actions are selected infinitely often.
Control: The resulting experience includes different actions and can support Monte Carlo control while the same policy remains the one being evaluated or improved.
On-policy Monte Carlo control avoids exploring starts by maintaining action opportunities over repeated decisions rather than requiring them all at the initial condition.
Mistakes in Policy Reasoning
Treating exploring starts as an ordinary guarantee of every environment.
Exploring starts is a convenient assumption, but the source identifies it as unrealistic because the required starting condition may not be available.
Fix:
Use continued action selection so that all actions are selected infinitely often.Thinking that avoiding exploring starts means avoiding different actions.
Monte Carlo control needs experience that includes different actions.
Fix:
Keep selecting actions across episodes so that all actions continue to receive selection opportunities.Calling a method on-policy merely because it uses episodes.
The on-policy and off-policy distinction concerns the relationship between those policies.
Fix:
Call the method on-policy when it evaluates or improves the same policy used to make decisions; call it off-policy when it evaluates or improves a different policy.
Check Your Understanding
Explain, in your own words, how an on-policy Monte Carlo control method can obtain experience with different actions without assuming that every action is explored at the beginning.
Hints
- Start by naming the limitation of exploring starts.
- Then describe what must happen across repeated action selections.
- Finally, identify which policy is evaluated or improved in the on-policy approach.
Key Takeaways
- Exploring starts assumes that all actions can be explored at the beginning, but that starting condition may be unrealistic. The general alternative is to continue selecting actions so that all actions are selected infinitely often. An on-policy method evaluates or improves the same policy used to make decisions, while an off-policy method evaluates or improves a different policy. Monte Carlo control without exploring starts follows the on-policy route and maintains action opportunities across repeated episodes instead of relying on an unrealistic initial condition.
Key Takeaways
- Exploring starts assumes that all actions can be explored at the beginning, an assumption that may not be available in a real setting.
- The general alternative is to keep selecting actions so that all actions are selected infinitely often.
- On-policy methods evaluate or improve the same policy that makes decisions.
- Off-policy methods evaluate or improve a different policy from the one that generates the data.
- On-policy Monte Carlo control avoids exploring starts by maintaining action selection opportunities across repeated episodes.