Concepts / Searching the Space of Policies

Searching the Space of Policies

Evolutionary methods evaluate a fixed policy through many game outcomes.

  • Programming

One Policy Under Test

When an agent plays a game, a policy determines how it behaves. Searching the space of policies means considering different candidate ways for the agent to act and using evaluation information to guide which policy should be considered later. Two important approaches organize that evaluation differently: evolutionary methods judge a fixed policy from many complete game outcomes, while value function methods evaluate individual states during play.

Imagine evaluating a Tic-Tac-Toe policy without changing it while you test it. The evolutionary method holds that policy fixed and plays many games against an opponent. The policy is judged from the outcomes of those games rather than from an evaluation of each individual position encountered during play.

is tested inproduceare combined intoFixed policyunchanged during testingMany gamesagainst an opponentGame outcomesresults of the gamesWin frequencypolicy evaluation
How does one fixed policy produce many game outcomes that are combined into an evaluation?

From Games to Win Frequency

The central trace for an evolutionary method is policy to games, games to final outcomes, and outcomes to win frequency. The policy remains fixed while many games are played. Instead of assigning credit by examining every position encountered during those games, the method uses the collection of final outcomes to judge the policy as a whole.

Win frequency is the frequency with which a fixed policy wins across the games used for evaluation. It estimates the probability of winning and guides the selection of a later policy.

Comparing Two Candidate Policies

Two candidate policies are tested through games against an opponent. How can their win frequencies guide a later policy selection?

Hold each policy fixed: Test one candidate without changing it during its games, then test the other candidate in the same broad way.

Collect final outcomes: For each candidate, use the outcomes of its games rather than evaluating each individual position encountered during play.

Estimate win frequency: Use how often each candidate wins to estimate its probability of winning.

Guide the later choice: Use the win-frequency estimates as evidence when selecting a later policy. A candidate with stronger evidence from its game outcomes provides a stronger basis for that later selection.

The policy is evaluated at the level of many game outcomes, and the resulting win frequency guides which policy is considered later.

are evaluated throughproduceguideCandidate policiespolicies underconsiderationMany gamesfor each candidateWin frequenciesestimated winningprobabilitiesLater policyselectionguided by the estimates
How do the win frequencies of candidate policies determine which policy is selected or reproduced next?

Evaluating States During Play

A value function method follows a different trace. During play, it allows individual states to be evaluated. A state is a position reached during the game, and the method uses information that appears during the game instead of waiting for the evaluation to be represented only by the final result of a collection of games.

is evaluated assupports information duringis evaluated asGame stateposition during playState evaluationvalue estimateLater game stateanother position duringplayLater evaluationinformation availableduring play
How does a value function assign an estimate to each game state as play progresses?

In the Tic-Tac-Toe setting, the evolutionary method waits for the outcomes of many games to judge the fixed policy. A value function method instead can evaluate the individual positions encountered as those games are being played. The difference is not simply the number of games; it is the point at which useful evaluation information is available.

Two Routes Through Policy Space

QuestionEvolutionary methodsValue function methods
What is evaluated?A fixed policy through many game outcomesIndividual states during play
When is information used?After outcomes from many games are collectedAs information appears during the game
What information receives attention?Final game outcomes and their win frequencyEvaluations of positions encountered during play
How is policy progress guided?Win frequency guides selection of a later policyState evaluations provide information while play is continuing

Both approaches can be understood as searching through possible policies, but they use different evidence while doing so. Evolutionary methods attach their evaluation to a policy as a whole: keep the policy fixed, observe many outcomes, estimate its win frequency, and use that estimate to guide a later policy selection. Value function methods attach evaluation to individual states encountered during play, so information from inside the games contributes while play is unfolding.

usesguidesevaluatesprovidesEvolutionary methodpolicy-level searchGame outcomeswin frequencyLater policyselection guided byoutcomesValue functionmethodstate-level searchIndividual statesevaluated during playInformation duringplayavailable before gamecollection ends
What information does each method use while searching among candidate policies?
can be searched bysummarizes asguidescan be searched byusesinformsPolicy spacecandidate policiesEvolutionaryevaluationmany game outcomesWin frequencyguides later selectionValue evaluationindividual statesInformation duringplaystate-based evidenceLater policy behaviorsearch informed byevaluation
How do evolutionary and value function methods move through different candidate policies toward better behavior?

Common Reasoning Errors

  • Treating an evolutionary evaluation as an evaluation of every position in a game.

    The evolutionary method judges a fixed policy from the outcomes of many games rather than from an evaluation of each individual position.

    Fix: Trace the evaluation from the fixed policy to many games, then to final outcomes, and finally to win frequency.

  • Assuming that one game is enough to establish a policy's evaluation.

    The method evaluates the fixed policy through many game outcomes.

    Fix: Think in terms of a collection of games and the win frequency estimated from their outcomes.

  • Confusing win frequency with a description of a particular game state.

    Win frequency summarizes outcomes for the policy, whereas value function methods evaluate individual states during play.

    Fix: Keep the evaluation units separate: policy-level outcome frequency versus state-level evaluation.

  • Claiming that the evolutionary method changes the policy after every move.

    The source describes the policy as fixed during testing and says that a policy change comes only after many games.

    Fix: Place the policy change after the collection of game outcomes has been evaluated.

When comparing the methods, ask two questions: What is being evaluated, and when does the evaluation information become available? For evolutionary methods, the answers are a fixed policy and the outcomes of many games. For value function methods, the answer is individual states and information available during play.

Check Your Understanding

EASY

A policy is held fixed while it plays many games against an opponent. The evaluator records the final outcomes and computes a win frequency. Is this evaluation policy-level or state-level, and which method does it describe? Then contrast it with the information used by a value function method.

Hints
  • Identify whether the evaluator is looking at complete game outcomes or individual positions.
  • Remember that the fixed-policy evaluation uses win frequency to guide a later policy selection.
  • For the contrast, focus on information available during play.

What do you think happens?

A method evaluates a fixed policy through many games and uses the resulting win frequency to guide a later policy selection. What happens if you describe the method as evaluating each individual position during play?

  • That description matches the evolutionary method exactly.
  • That description confuses the evolutionary method with the value function approach.
  • That description says the policy is never evaluated.
Reveal answer

Answer: That description confuses the evolutionary method with the value function approach.

The evolutionary method judges the fixed policy from many game outcomes. The value function approach is the one described as evaluating individual states during play.

The Essential Distinction

  1. Evolutionary methods keep a policy fixed while testing it through many games.
  2. The final outcomes of those games produce a win frequency that estimates the probability of winning.
  3. Win frequency guides the selection of a later policy.
  4. Value function methods evaluate individual states during play and use information available as the game unfolds.
  5. Both approaches search among policies, but evolutionary methods emphasize policy-level game outcomes while value function methods emphasize state-level information during play.

Key Takeaways

  • A fixed policy can be evaluated by playing many games and combining their outcomes into a win frequency.
  • Win frequency estimates the probability of winning and guides a later policy selection.
  • Evolutionary methods judge a policy broadly from final outcomes rather than evaluating every position encountered during play.
  • Value function methods evaluate individual states while play is occurring.
  • The approaches search policy space using different information: collections of game outcomes versus state evaluations available during play.