Concepts / Evolutionary Methods

Evolutionary Methods

A policy gradient method searches over parameter-defined policies rather than treating a value estimate as the primary search object.

  • Machine Learning

Two Ways to Search for Better Behavior

Reinforcement learning methods can search for better behavior from different starting points. A policy gradient method searches through policies defined by numerical parameters. It uses interaction with the environment to estimate a direction for changing those parameters. An evolutionary method takes another route: it compares the complete behavior of several candidate policies and favors candidates that obtain more reward over their interaction period.

The central question is not simply whether a method uses reward. Both approaches use reward-related evidence. The key difference is what the method treats as the main object of search and evaluation.

Searching Through Policy Parameters

A policy gradient method searches over parameter-defined policies rather than treating a value estimate as the primary search object. The policy is described by a collection of numerical parameters. Those parameters determine the policy the agent uses, so changing the parameters changes the policy itself.

defineestimatecomparePolicy parametersNumerical parametersValue estimateState or action valueCandidate policiesComplete behaviorsDefined policyPrimary search object
What is being searched over in a policy gradient method, and how does that differ from estimating values or evolving complete policies?

The diagram separates three ideas that can otherwise be confused. Policy gradient search changes the numerical parameters that define a policy. Value-function estimation focuses on how valuable a state or action is. Evolutionary search compares candidate policies by the behavior those policies produce over an interaction period. These ideas can appear together in reinforcement learning, but they are not identical.

From Interaction to Parameter Adjustment

defineacts inproducesinformsadjustsPolicy parametersCurrent numerical settingsPolicyParameter-defined behaviorEnvironmentInteractionExperienceBehavioral interactiondetailsImprovement directionEstimated parameter changeAdjusted parametersNew numerical settings
How does data move from policy-environment interaction into a performance measurement or a policy-parameter update?

Environment interaction matters because the method needs evidence about how the current policy behaves. The policy uses its current numerical parameters, the agent interacts with the environment, and the resulting behavioral interactions provide information for estimating a direction in which the parameters should move. The method then adjusts the parameters and searches again with another parameter-defined policy.

definedefineParameters ACurrent valuesParameters BAdjusted valuesPolicy ABehavior from Parameters APolicy BBehavior from Parameters B
What changes after an estimated improvement direction is applied to the policy parameters?

Tracing One Policy Search Step

Suppose a policy is defined by numerical parameters and produces an interaction with the environment. What does a policy gradient method do with that interaction?

Start with a policy: The current numerical parameters define the policy the agent uses.

Interact with the environment: The agent follows the policy, and the resulting behavioral interactions provide experience.

Estimate a direction: The experience is used to estimate a direction for changing the policy parameters so the policy may perform better.

Adjust and search again: The parameters are changed in the estimated direction, producing another parameter-defined policy.

The method searches by repeatedly using environment interaction to guide changes in the numerical parameters that define the policy.

The Role of Value Estimates

Value functions are not required in every method in this family. Some policy gradient methods use value-function estimates to improve their gradient estimates. In that situation, the value estimate is not the same thing as the policy-gradient estimate. It helps make the estimated direction for changing policy parameters more useful.

can informcan informcan improveguides adjustmentEnvironmentexperienceAgent interactionValue estimateState or actionGradient estimateDirection for parameterchangePolicy parametersAdjusted search object
How can a state-value or action-value estimate alter the signal used to estimate which policy-parameter changes may improve performance?

Evaluating a Population of Policies

Evolutionary methods approach reinforcement learning by comparing complete behaviors rather than constructing a value estimate. Imagine a collection of agents that use different policies. Each agent interacts with the environment during its lifetime. After that interaction period, the agent is evaluated by the reward obtained from its complete behavior. Candidates with more reward are favored as the basis for further search, while weaker candidates are not favored for the next stage.

usesusesusesproducesproducesproducescomparedcomparedcomparedbasis forCandidate policiesDifferent policy choicesAgent ALifetime behaviorReward AComplete behaviorSelected candidatesHigher-reward basisFurther searchNext stageAgent BLifetime behaviorReward BComplete behaviorAgent CLifetime behaviorReward CComplete behavior
How are multiple policy agents evaluated in the environment, compared by lifetime performance, and selected for the next stage of the search?

Comparing Complete Behaviors

Three non-learning agents use different policies during an interaction period. Agent A obtains more reward than Agents B and C. How does the evolutionary method use this result?

Evaluate each lifetime: Each agent is judged by the reward obtained from its complete interaction behavior.

Compare the results: The rewards associated with the candidates are compared against one another.

Favor the stronger candidate: Because Agent A obtained more reward, it is favored as a basis for further search.

Do not require lifetime learning: The individual agents were described as non-learning agents. The broader search favors candidates after their behavior is evaluated.

The evolutionary process selects according to comparative lifetime performance, not according to a value estimate attached to one isolated state or action.

Lifetime Performance Versus Value

evaluated bycan be evaluatedcan be evaluatedLifetime behaviorComplete interaction periodStateParticular situationReward resultOne comparative scoreActionParticular choiceValue estimateState or action value
What is the difference between assigning one fitness score to an agent's whole lifetime behavior and estimating expected return from a particular state or action?
Evaluation viewpointWhat is evaluatedRole in the search
Evolutionary methodAn entire policy's behavior over its interaction periodCandidates with more reward are favored for further search
Value-function estimateA state or an actionDescribes how valuable that state or action is
Policy gradient methodA policy defined by numerical parametersEstimates a direction for changing those parameters

An evolutionary method can compare whole policies even when it never constructs a value function. Its important unit of evaluation is the agent's lifetime behavior. A value function instead estimates the value of a state or action. These approaches are different, although both can be used in reinforcement learning and some policy gradient methods can use value estimates to improve their gradient estimates.

When Evolutionary Search Helps

Evolutionary methods may be effective when the policy space is small, when the policy space can be organized so that good policies are common or easy to find, or when substantial time is available for searching. These conditions make it more plausible that repeated comparison and selection will encounter useful policies.

Common Reasoning Errors

  • Saying that a policy gradient method primarily searches over value estimates.

    The primary search object is a policy defined by numerical parameters.

    Fix: State that the method estimates a direction for adjusting the parameters that define the policy.

  • Claiming that every policy gradient method must use a value function.

    Value estimates can improve gradient estimates in some methods, but they are not required in all of them.

    Fix: Describe value estimates as optional aids in some policy gradient methods.

  • Treating an evolutionary method's score as a value function.

    The evolutionary method evaluates complete lifetime behavior, whereas a value function estimates the value of a state or action.

    Fix: Call the lifetime result a reward-based comparison of candidate policies.

  • Assuming each agent must learn during its lifetime for the method to be evolutionary.

    The source describes non-learning agents whose complete behaviors are evaluated, followed by selection in the broader search.

    Fix: Separate individual lifetime behavior from the broader process that favors stronger candidates.

Practice the Distinction

MEDIUM

A method changes the numerical parameters of one policy after using experience from interaction with the environment. Another method evaluates several complete policies, compares the rewards obtained over their interaction periods, and favors candidates with better results. Identify the main search object and evaluation viewpoint for each method.

Hints
  • Ask what is directly changed in the first method.
  • Ask whether the second method evaluates one state or action, or an entire policy's behavior.
  • Remember that a value estimate may support some policy gradient methods without being required by all of them.

What do you think happens?

A candidate policy receives a high reward over its complete interaction period, but the method never estimates the value of an individual state or action. Can this still be an evolutionary method?

  • Yes
  • No
Reveal answer

Answer: Yes

Evolutionary methods can compare complete lifetime behaviors by their reward without constructing a state-value or action-value estimate.

Key Takeaways

  1. Policy gradient methods search through policies defined by numerical parameters.
  2. Environment interaction provides evidence for estimating a direction in which those parameters should be adjusted.
  3. Value-function estimates can improve gradient estimates in some policy gradient methods, but they are not required in every method.
  4. Evolutionary methods compare complete policy behaviors over an interaction period and favor candidates that obtain more reward.
  5. Evolutionary evaluation differs from estimating the value of a particular state or action and may be effective in small or searchable policy spaces, with substantial search time, or when accurate environmental sensing is difficult.

Key Takeaways

  • Policy gradient search changes the numerical parameters that define a policy.
  • Interaction with the environment supplies the evidence used to estimate a useful parameter-adjustment direction.
  • Value estimates may support gradient estimation, but they are not required by every policy gradient method.
  • Evolutionary methods evaluate complete lifetime behaviors, compare their rewards, and favor stronger candidate policies.
  • The boundary between methods is not absolute, because value estimates and policy-gradient estimates can appear together, but their primary search and evaluation viewpoints remain distinct.