Monte Carlo Policy Gradient Methods
This practice section applies REINFORCE to gridworlds.
From Gridworld Practice to Results
This section uses REINFORCE in the gridworlds from Examples 13.1 and/or 13.2. The gridworld is not introduced as a separate topic here. It is the setting in which the algorithm is applied, so the reported results belong to that application.
When reviewing a REINFORCE result, keep two questions connected: which gridworld example was used, and what result did applying the algorithm produce in that setting? The practice section establishes this application and its connection to the reported results. It does not, by itself, establish every internal detail of the algorithm.
Tracing the Policy Search
A policy gradient method searches over policies defined by numerical parameters rather than treating a value estimate as the primary search object.
The method begins with numerical parameters. Those parameters define the policy used by the agent. The agent then interacts with the environment, and the resulting experience is used to estimate a direction for changing the parameters. After the adjustment, the parameters define another policy. Repeating this process makes the search a movement through possible parameter-defined policies.
The central estimate is therefore a direction for changing the policy's numerical parameters. It is not primarily a search for a value estimate. Value estimates may participate in some methods, but the object being directly searched is the parameter-defined policy.
What Changes After Interaction
A Conceptual Gridworld Trace
Trace what changes when REINFORCE is applied in one of the stated gridworld settings.
1. Start with a policy: The agent begins with numerical parameters that define the policy it will use in the gridworld.
2. Interact with the environment: The policy guides the agent's behavior, and the interaction produces experience in the gridworld.
3. Use the resulting returns: The experience provides information used to estimate a direction for improving the policy parameters.
4. Adjust the parameters: The numerical parameters are changed in the estimated improvement direction.
5. Interpret the new policy: After the adjustment, the parameters define another policy. The reported result is understood as the outcome of applying REINFORCE in the selected gridworld.
The meaningful change is from one parameter-defined policy to another, based on information obtained through interaction with the environment.
The interaction matters because the improvement direction is estimated from the agent's experience with the environment. Without that interaction, the method would not have the experience-based information used to decide how the policy parameters should move.
Value Estimates and Gradient Quality
Value functions are not required in every policy gradient method. Some methods do not appeal to value functions. However, some policy gradient methods use value function estimates to improve their gradient estimates. In that role, the value estimate helps make the estimated improvement direction more useful; it does not replace the policy gradient as the object used to adjust the policy parameters.
Boundaries with Evolutionary Search
| Aspect | Policy gradient perspective | Evolutionary perspective |
|---|---|---|
| Search object | Policies defined by numerical parameters | Search through evolutionary variations |
| Main improvement signal | An estimated direction for changing policy parameters | Population-based evolutionary search |
| Role of environment interaction | Experience from interaction estimates the direction | The source distinction is not absolute |
The practical distinction is that policy gradient methods are described as searching through parameter-defined policies and estimating how their parameters should move. Evolutionary methods are discussed as using evolutionary variations. These descriptions help distinguish the approaches, but they do not justify treating the boundary between them as absolute.
Mistakes in Reading the Practice Section
Treating the gridworld as the main definition of REINFORCE.
The gridworld is the setting used to apply REINFORCE in this practice section.
Fix:
Use the gridworlds from Examples 13.1 and/or 13.2 as the application context, then connect the application to its reported results.Saying that policy gradient methods search primarily over value estimates.
Policy gradient methods search over policies defined by numerical parameters.
Fix:
State that experience is used to estimate a direction for adjusting the policy parameters.Claiming that every policy gradient method requires a value function.
Some methods do not appeal to value functions.
Fix:
Explain that value function estimates can improve gradient estimates in some methods, but are not required in all of them.Describing the update as a switch between named algorithms.
The meaningful change is in the numerical parameters that define the policy.
Fix:
Describe the process as moving from one parameter-defined policy to another.
Practice Check
A learner says: REINFORCE is practised in a gridworld, so the practice section defines all of REINFORCE's internal mechanics. Correct the statement in two parts: first identify the setting, and then state what the section does and does not establish.
Hints
- Name Examples 13.1 and/or 13.2 as the relevant gridworld settings.
- Separate the practical application and its reported results from algorithm details that must be learned elsewhere.
What do you think happens?
Before reading the explanation, what changes after the agent's interaction is used to estimate an improvement direction?
Reveal answer
Answer: The numerical parameters defining the policy
The source describes the meaningful change as an adjustment to the numerical parameters that define the policy. The resulting parameters define another policy.
Key Takeaways
- This section applies REINFORCE to the gridworlds from Examples 13.1 and/or 13.2.
- The application connects the algorithm with the results reported for those gridworld settings.
- A policy gradient method searches through policies defined by numerical parameters.
- Interaction with the environment provides the experience used to estimate a direction for changing those parameters.
- Value function estimates can improve gradient estimates in some methods, but they are not required in every policy gradient method.
Key Takeaways
- REINFORCE is practised in the gridworlds from Examples 13.1 and/or 13.2.
- The practice section establishes an application and its connection to reported results, not every internal algorithm detail.
- Policy gradient methods search over parameter-defined policies and adjust their numerical parameters using experience from environment interaction.
- Value function estimates may improve gradient estimates in some methods, but they are not required by all policy gradient methods.
- The distinction between policy gradient and evolutionary methods is useful, but their boundary should not be treated as absolute.