Concepts / Reinforcement Learning Methods

Reinforcement Learning Methods

A policy gradient method searches over parameter-defined policies rather than treating a value estimate as the primary search object.

  • Machine Learning

From Parameters to Behavior

Many reinforcement learning discussions begin by asking how good a situation or action is. A policy gradient method begins from a different angle: it searches directly through policies. Each policy is described by a collection of numerical parameters. The method uses experience from interaction with the environment to estimate how those parameters should change so that the policy performs better.

The primary search object is a parameter-defined policy, not necessarily a value estimate.

defineselectsdefineselectsPolicy parametersbefore adjustmentPolicyparameter-defined behaviorActionsselected in statesPolicy parametersafter adjustmentPolicynew parameter-definedbehaviorActionsselected in states
What changes when an estimated improvement direction is applied to the numerical parameters that define a policy?

The Interaction Loop

A policy gradient method cannot decide how its parameters should move from the parameters alone. It needs information produced by the agent's interaction with the environment. The current parameters define the policy being used. The agent follows that policy, interacts with the environment, and obtains behavioral information from the resulting interaction. That information is used to estimate a direction for improving the policy. The parameters can then be adjusted, producing another parameter-defined policy.

defineacts inprovides context forproduces interactionestimatesadjustsdefines nextPolicy parametersnumerical parametersPolicydefined by parametersEnvironment statecurrent situationActionselected by policyEnvironment feedbackinteraction informationImprovement directionestimated parameter changeUpdated parametersnext policy search point
How do parameters, actions, environment feedback, and the next update connect across repeated interaction?

Improving the Gradient Estimate

The central estimate in a policy gradient method is a direction for changing the policy's numerical parameters. Some methods improve this estimate with a value-function estimate. The value estimate can make the estimated improvement direction more useful, but it is not the same thing as the policy gradient estimate.

provides information forinformsimprovesadjustsEnvironmentinteractionexperienceReturn or valueestimateestimated outcomeinformationPolicy improvementsignalbaseline or advantageinformationGradient estimatedirection for parameterchangePolicy parametersadjusted for another searchstep
How can a return or value estimate provide information that improves the estimated direction for changing policy parameters?

Do not treat the presence of a value estimate as the definition of every policy gradient method. The source distinguishes the two ideas: value estimates can improve gradient estimates in some methods, while other methods do not appeal to value functions.

Backup Dimensions

Reinforcement learning methods can be organized along several independent dimensions. First ask how information is obtained: a sample backup uses a sampled trajectory, whereas a full backup uses a distribution of possible trajectories and requires a model. Then ask how far the backup looks ahead. Finally ask how the value function is represented. These questions should remain separate because one choice does not determine the others.

requirescan usecan usecan usecan usecan useSample backupone observed trajectoryFull backuppossible trajectoriesOne-step TDstrong bootstrappingn-step methodintermediate depthFull-return MonteCarlolittle or no bootstrappingfrom a later estimateModelrequired by full backups
How do sample versus full backups differ in breadth, and how does one-step versus full-return backup relate to bootstrapping?

Backup kind describes breadth, not distance. A sample backup may use one observed path, while a full backup may account for possible paths. Backup depth is a separate question. It ranges from one-step temporal-difference updates, through intermediate n-step and mixed methods, to full-return Monte Carlo updates.

Backup depth also describes the amount of bootstrapping. One-step temporal-difference backups rely strongly on an estimated value after the immediate step. Full-return Monte Carlo backups look through the complete return instead. Intermediate methods occupy the space between these choices.

Representation Spectrum

A third dimension concerns how the value function is represented. The spectrum runs from tabular representation through state aggregation and linear methods to nonlinear methods. This representation choice is independent of backup kind and backup depth. For example, two methods can use the same backup idea while representing the value function in different ways.

classify byclassify byclassify byranges fromthroughthroughtoInformation sourcesample or full backupBackup depthone-step to full returnValuerepresentationtabular to nonlinearTabularrepresentation endpointState aggregationintermediate representationLinearfunction approximationNonlinearrepresentation endpointReinforcementlearning methodlocated on all threedimensions
Where should a reinforcement learning method be placed when comparing information source, backup depth, and value-function representation?

Search Extremes

The depth dimension also helps position search methods. At one extreme, one-step temporal-difference backups use an estimated value after the immediate step. At the other extreme, exhaustive search carries full backups to terminal states. In a continuing task, search can continue until discounting makes later rewards contribute negligibly. Search methods between these extremes may look ahead only to a limited depth and may select which possibilities to explore.

increasing depthincreasing depthcan continue until later contributions are negligibleOne-step TDestimated value after onestepIntermediate searchlimited lookaheadExhaustive searchfull backups to terminalstatesContinuing tasklater rewards discounted
How does one-step bootstrapping compare with exhaustive search that carries backups through a complete return?

Classifying Hypothetical Methods

Separating Backup Kind from Backup Depth

Compare two hypothetical value-learning methods. Method A updates from one observed trajectory and stops its backup after one step. Method B uses a model to consider possible trajectories and continues its backup through the full return.

Method A information source: Method A uses one observed trajectory, so it uses a sample backup.

Method A depth: Method A stops after one step, placing it at the one-step temporal-difference end of the backup-depth dimension.

Method B information source: Method B uses a model to consider possible trajectories, so it uses a full backup.

Method B depth: Method B continues through the full return, placing it at the full-return end rather than the one-step end.

Representation check: The example demonstrates information source and depth. A value-function representation should be classified separately; neither method's representation is specified here.

Method A is a sample, one-step method. Method B is a full-backup, full-return method. The example shows why backup kind and backup depth must not be treated as the same dimension.

QuestionMethod AMethod B
Where does information come from?One observed trajectoryA model and possible trajectories
Backup kindSample backupFull backup
How far does the backup go?One stepFull return
Bootstrapping positionStrongly associated with one-step estimated value informationLooks through the complete return
Value representationNot specifiedNot specified

Common Classification Mistakes

  • Treating a value function as required by every policy gradient method.

    The source states that some policy gradient methods do not appeal to value functions.

    Fix: Say that value-function estimates can improve gradient estimates in some methods, but are not required in all methods.

  • Confusing the policy with the value estimate.

    The primary search object is a policy defined by numerical parameters.

    Fix: Track the policy parameters and the direction estimated for changing them.

  • Treating sample backups and full backups as different backup depths.

    Backup kind concerns breadth of information, while backup depth concerns how far ahead the backup looks.

    Fix: Classify information source and lookahead depth separately.

  • Assuming one-step and full-return backups are the only possible depths.

    The source describes intermediate methods between one-step temporal-difference and full-return Monte Carlo updates.

    Fix: Place intermediate methods between the two endpoints.

  • Using one label to describe every dimension of a method.

    Methods can differ independently in information source, lookahead depth, and value-function representation.

    Fix: Use the three-pass procedure: information source, depth, representation.

Practice: Three-Pass Classification

MEDIUM

A hypothetical method learns from one sampled trajectory, uses an intermediate n-step backup, and represents its value function with a linear method. Classify the method along the three independent dimensions.

Hints
  • First identify whether the information comes from a sampled trajectory or a modelled distribution of possible trajectories.
  • Next place the backup depth between one-step temporal-difference and full-return Monte Carlo.
  • Finally locate the representation on the spectrum from tabular through state aggregation and linear methods to nonlinear methods.
EASY

Explain why a policy gradient method still needs environment interaction even though its search object is a set of policy parameters.

Hints
  • The parameters define the policy but do not by themselves provide the improvement direction.
  • The direction is estimated from the agent's interaction with the environment.

Method Comparison Checklist

  1. Policy gradient methods search through policies defined by numerical parameters.
  2. Environment interaction provides the information used to estimate a direction for changing those parameters.
  3. Value-function estimates can improve gradient estimates in some methods, but they are not required in every policy gradient method.
  4. Sample backups use sampled trajectories; full backups use a distribution of possible trajectories and require a model.
  5. Backup depth ranges from one-step temporal-difference updates through intermediate methods to full-return Monte Carlo updates, with bootstrapping decreasing as the backup reaches further into the return.
  6. Value-function representation ranges from tabular methods through state aggregation and linear methods to nonlinear methods, so methods should be compared across independent dimensions.

Key Takeaways

  • A policy gradient method searches over policies defined by numerical parameters and estimates how those parameters should change.
  • Interaction with the environment supplies the experience needed to estimate an improvement direction.
  • Value estimates may improve gradient estimates, but value functions are not required by every policy gradient method.
  • Sample versus full backup describes information breadth, while one-step through full-return backup describes depth and bootstrapping.
  • Compare reinforcement learning methods separately by information source, backup depth, and value-function representation.