Concepts / Policy Search in Reinforcement Learning

Policy Search in Reinforcement Learning

Evolutionary methods use policy search and reward comparison rather than value-function estimation.

  • Machine Learning

A Different Route to Better Behavior

Many reinforcement learning methods try to estimate how valuable particular states or actions are. Evolutionary methods take a different route: they search among policies by comparing the rewards produced by complete episodes or interaction periods. A candidate is judged by the behavior its policy produces over its lifetime, and candidates associated with more reward are favored for further search.

The central shift is from estimating the value of individual states or actions to comparing the outcome of an entire policy.

One Generation of Policy Search

use policiesproducescomparefavorcontinue searchCandidate policiesinitial groupEnvironmentinteractioneach lifetimeLifetime rewardbehavior outcomeSelected candidatesmore reward favoredFurther searchnext candidate group
How are policies initialized, run through interaction periods, evaluated by reward, selected, and varied to produce the next generation?

A policy-search method begins with a collection of candidate policies. Each candidate controls an agent during an interaction period. The agent does not improve through learning during its own lifetime. When the interaction period ends, the method evaluates the reward produced by that complete behavior. Candidates with better results are favored as the basis for further search.

Comparing Three Candidate Policies

Suppose three non-learning agents use different policies during the same type of interaction period. Determine which candidates are favored after their behavior is evaluated.

Run each policy: Each candidate agent follows its own policy while interacting with the environment. No candidate improves through learning during that individual lifetime.

Evaluate complete behavior: After the interaction period, evaluate the reward obtained by each agent over that lifetime.

Compare outcomes: Place the candidates in order according to the reward their complete behaviors produced.

Continue the search: Candidates associated with more reward are favored as the basis for further search, while weaker candidates are not favored for the next stage.

The method selects according to lifetime reward, not according to a separately estimated value for one state or one action.

Comparing a Population

lifetime rewardlifetime rewardlifetime rewardhigher rewardlower rewardPolicy Aagent lifetimeReward comparisoncomplete behaviorsFavored policiesmore rewardPolicy Bagent lifetimeWeaker candidatesnot favoredPolicy Cagent lifetime
How are multiple agents' lifetime rewards compared, and which policies continue into the next stage of the search?

The population is important because selection is comparative. A policy is not judged by inspecting only one isolated decision. Each candidate is run, its complete interaction behavior is evaluated, and its result is compared with the results of other candidates. The candidates with better reward become the stronger basis for the next stage of the search.

Lifetime Results and Value Estimates

evaluateestimatePolicycomplete behaviorLifetime rewardone behavior outcomeState or actionindividual choice pointValue estimateestimated value
What is the difference between assigning one total reward to a policy's interaction period and estimating values for individual states or actions?
Question being askedPolicy-search viewpointValue-estimation viewpoint
What is evaluated?The behavior produced by an entire policy over its interaction periodThe value of a state or action
What is compared or estimated?Reward obtained by complete candidate behaviorsEstimated value attached to individual states or actions
What does the method need to construct?A comparison-and-selection process over policiesA value estimate
Can one approach exist without the other?Yes; evolutionary methods do not require a value functionValue-based methods follow the value-estimation route

A value function estimates how valuable a state or action is. Policy search instead asks how well an entire policy performed over its interaction period. This difference matters because the unit of evaluation is the agent's lifetime behavior, not an isolated estimate attached to one state or one action. The two approaches can both address reinforcement learning problems, but policy search does not require first constructing a value function.

When Search May Work Well

Evolutionary methods may be effective when the policy space is small, when the policy space is organized so that good policies are common or easy to find, or when substantial time is available for searching. These conditions make it more plausible that repeated comparison and selection will encounter useful policies.

  • A small policy space can make the search manageable.
  • A policy space containing many good or easy-to-find policies can make reward-based selection more promising.
  • Substantial available search time can allow the comparison-and-selection process to encounter useful policies.
  • Limited ability to sense the environment can favor this approach when methods depending on accurate environmental sensing would face difficulty.

Mistakes in Choosing the Evaluation Unit

  • Treating policy search as if it estimates the value of every state or action.

    The defining contrast is that evolutionary methods compare complete behaviors by the reward obtained over an interaction period rather than building a value estimate.

    Fix: Describe the candidate policy, its lifetime behavior, its resulting reward, and its comparison with other candidates.

  • Assuming that an individual candidate learns during its own lifetime.

    The source describes non-learning agents whose policies are evaluated after the interaction period.

    Fix: Separate individual behavior from broader search: the individual follows its policy, while selection across candidates favors better results.

  • Evaluating only one isolated action instead of the complete behavior.

    The important unit of evaluation is the agent's lifetime behavior and the reward obtained from that complete behavior.

    Fix: Compare the reward produced over the candidate's interaction period.

  • Claiming that evolutionary methods are always preferable.

    The source identifies conditions that may make evolutionary methods effective, such as a small or searchable policy space, available search time, or limited environmental sensing.

    Fix: Treat effectiveness as dependent on the characteristics of the policy space, search budget, and sensing situation.

Trace the Selection Process

MEDIUM

Explain the following process in your own words: several agents use different policies, interact with the environment, receive rewards for their complete behaviors, and are compared at the end of the interaction period. Your explanation must identify which candidates are favored and state why this process does not require a value function.

Hints
  • Name the unit that is evaluated.
  • Distinguish the reward obtained by a complete behavior from an estimate attached to one state or action.
  • Explain what happens to candidates associated with more reward.

What do you think happens?

Before reading the explanation, predict which candidate is favored: one candidate has a strong result over its complete interaction period, while another has one apparently useful action but a weaker overall result.

  • The candidate with the stronger complete-interaction result
  • The candidate with the single apparently useful action
  • Neither candidate, because policy search requires a value estimate first
Reveal answer

Answer: The candidate with the stronger complete-interaction result

Evolutionary evaluation compares complete policy behaviors by the reward obtained over the interaction period. It does not select from an isolated action and does not require a value function.

The Search View in One Pass

  1. Evolutionary reinforcement learning searches among policies through reward comparison rather than value-function estimation.
  2. Each candidate is judged by the behavior produced over its lifetime or interaction period.
  3. Candidates associated with more reward are favored as the basis for further search.
  4. The method evaluates complete behaviors, whereas a value function estimates the value of a state or action.
  5. Evolutionary methods may be effective with small or searchable policy spaces, substantial search time, or limited environmental sensing.

Key Takeaways

  • Policy search evaluates whole policies through the rewards produced by their complete interaction behavior.
  • Selection compares candidates in a population and favors those associated with more reward.
  • This evaluation differs from estimating the value of individual states or actions.
  • An individual candidate need not learn during its own lifetime; the broader search favors better-performing candidates.
  • Small or searchable policy spaces, available search time, and limited environmental sensing can make evolutionary methods effective.