Reinforcement Learning Methods
A policy gradient method searches over parameter-defined policies rather than treating a value estimate as the primary search object.
From Parameters to Behavior
Many reinforcement learning discussions begin by asking how good a situation or action is. A policy gradient method begins from a different angle: it searches directly through policies. Each policy is described by a collection of numerical parameters. The method uses experience from interaction with the environment to estimate how those parameters should change so that the policy performs better.
The primary search object is a parameter-defined policy, not necessarily a value estimate.
The Interaction Loop
A policy gradient method cannot decide how its parameters should move from the parameters alone. It needs information produced by the agent's interaction with the environment. The current parameters define the policy being used. The agent follows that policy, interacts with the environment, and obtains behavioral information from the resulting interaction. That information is used to estimate a direction for improving the policy. The parameters can then be adjusted, producing another parameter-defined policy.
Improving the Gradient Estimate
The central estimate in a policy gradient method is a direction for changing the policy's numerical parameters. Some methods improve this estimate with a value-function estimate. The value estimate can make the estimated improvement direction more useful, but it is not the same thing as the policy gradient estimate.
Do not treat the presence of a value estimate as the definition of every policy gradient method. The source distinguishes the two ideas: value estimates can improve gradient estimates in some methods, while other methods do not appeal to value functions.
Backup Dimensions
Reinforcement learning methods can be organized along several independent dimensions. First ask how information is obtained: a sample backup uses a sampled trajectory, whereas a full backup uses a distribution of possible trajectories and requires a model. Then ask how far the backup looks ahead. Finally ask how the value function is represented. These questions should remain separate because one choice does not determine the others.
Backup kind describes breadth, not distance. A sample backup may use one observed path, while a full backup may account for possible paths. Backup depth is a separate question. It ranges from one-step temporal-difference updates, through intermediate n-step and mixed methods, to full-return Monte Carlo updates.
Backup depth also describes the amount of bootstrapping. One-step temporal-difference backups rely strongly on an estimated value after the immediate step. Full-return Monte Carlo backups look through the complete return instead. Intermediate methods occupy the space between these choices.
Representation Spectrum
A third dimension concerns how the value function is represented. The spectrum runs from tabular representation through state aggregation and linear methods to nonlinear methods. This representation choice is independent of backup kind and backup depth. For example, two methods can use the same backup idea while representing the value function in different ways.
Search Extremes
The depth dimension also helps position search methods. At one extreme, one-step temporal-difference backups use an estimated value after the immediate step. At the other extreme, exhaustive search carries full backups to terminal states. In a continuing task, search can continue until discounting makes later rewards contribute negligibly. Search methods between these extremes may look ahead only to a limited depth and may select which possibilities to explore.
Classifying Hypothetical Methods
Separating Backup Kind from Backup Depth
Compare two hypothetical value-learning methods. Method A updates from one observed trajectory and stops its backup after one step. Method B uses a model to consider possible trajectories and continues its backup through the full return.
Method A information source: Method A uses one observed trajectory, so it uses a sample backup.
Method A depth: Method A stops after one step, placing it at the one-step temporal-difference end of the backup-depth dimension.
Method B information source: Method B uses a model to consider possible trajectories, so it uses a full backup.
Method B depth: Method B continues through the full return, placing it at the full-return end rather than the one-step end.
Representation check: The example demonstrates information source and depth. A value-function representation should be classified separately; neither method's representation is specified here.
Method A is a sample, one-step method. Method B is a full-backup, full-return method. The example shows why backup kind and backup depth must not be treated as the same dimension.
| Question | Method A | Method B |
|---|---|---|
| Where does information come from? | One observed trajectory | A model and possible trajectories |
| Backup kind | Sample backup | Full backup |
| How far does the backup go? | One step | Full return |
| Bootstrapping position | Strongly associated with one-step estimated value information | Looks through the complete return |
| Value representation | Not specified | Not specified |
Common Classification Mistakes
Treating a value function as required by every policy gradient method.
The source states that some policy gradient methods do not appeal to value functions.
Fix:
Say that value-function estimates can improve gradient estimates in some methods, but are not required in all methods.Confusing the policy with the value estimate.
The primary search object is a policy defined by numerical parameters.
Fix:
Track the policy parameters and the direction estimated for changing them.Treating sample backups and full backups as different backup depths.
Backup kind concerns breadth of information, while backup depth concerns how far ahead the backup looks.
Fix:
Classify information source and lookahead depth separately.Assuming one-step and full-return backups are the only possible depths.
The source describes intermediate methods between one-step temporal-difference and full-return Monte Carlo updates.
Fix:
Place intermediate methods between the two endpoints.Using one label to describe every dimension of a method.
Methods can differ independently in information source, lookahead depth, and value-function representation.
Fix:
Use the three-pass procedure: information source, depth, representation.
Practice: Three-Pass Classification
A hypothetical method learns from one sampled trajectory, uses an intermediate n-step backup, and represents its value function with a linear method. Classify the method along the three independent dimensions.
Hints
- First identify whether the information comes from a sampled trajectory or a modelled distribution of possible trajectories.
- Next place the backup depth between one-step temporal-difference and full-return Monte Carlo.
- Finally locate the representation on the spectrum from tabular through state aggregation and linear methods to nonlinear methods.
Explain why a policy gradient method still needs environment interaction even though its search object is a set of policy parameters.
Hints
- The parameters define the policy but do not by themselves provide the improvement direction.
- The direction is estimated from the agent's interaction with the environment.
Method Comparison Checklist
- Policy gradient methods search through policies defined by numerical parameters.
- Environment interaction provides the information used to estimate a direction for changing those parameters.
- Value-function estimates can improve gradient estimates in some methods, but they are not required in every policy gradient method.
- Sample backups use sampled trajectories; full backups use a distribution of possible trajectories and require a model.
- Backup depth ranges from one-step temporal-difference updates through intermediate methods to full-return Monte Carlo updates, with bootstrapping decreasing as the backup reaches further into the return.
- Value-function representation ranges from tabular methods through state aggregation and linear methods to nonlinear methods, so methods should be compared across independent dimensions.
Key Takeaways
- A policy gradient method searches over policies defined by numerical parameters and estimates how those parameters should change.
- Interaction with the environment supplies the experience needed to estimate an improvement direction.
- Value estimates may improve gradient estimates, but value functions are not required by every policy gradient method.
- Sample versus full backup describes information breadth, while one-step through full-return backup describes depth and bootstrapping.
- Compare reinforcement learning methods separately by information source, backup depth, and value-function representation.