Concepts / AlphaGo's Rollout Policy

AlphaGo's Rollout Policy

AlphaGo evaluated game states with both a value network and rollouts.

  • Programming

Two Views of One Go Position

When AlphaGo evaluated a Go game state, it did not rely on just one kind of judgment. It used a value network and rollouts as complementary evaluation methods. The value network evaluated the high-performance policy pρ, while rollouts used the weaker but faster policy pπ. The central design question was how much influence each evaluation should have.

evaluatesimulatecontributescontributesGo game stateValue networkEvaluates pρGame-state evaluationMixed using λRolloutsUses pπ
How do the value network and rollouts evaluate the same Go game state before their results are combined?

The two paths answer the same broad question: how should this Go position be evaluated? They do so through different mechanisms. The value network provides the learned evaluation associated with the high-performance policy pρ. Rollouts provide an evaluation based on the weaker but faster policy pπ. AlphaGo's rollout policy was therefore not a replacement for the value network. It was a second source of information about the same position.

Following λ Through the Evaluation

AlphaGo used λ as a mixing control. It determined how much the final game-state evaluation relied on the value network and how much it relied on rollouts. λ was not a third evaluation method and did not produce an independent judgment about the Go position. Its role was to control the balance between the two existing judgments.

one inputone inputcontrols mixValue networkInfluence decreases as λrisesFinal evaluationCombined resultλMixing controlRolloutsInfluence increases as λrises
What changes in the final evaluation as λ shifts the balance between the value-network estimate and the rollout estimate?

To interpret λ, ask one question: how much of the final evaluation is being assigned to rollouts rather than to the value network? A smaller λ gives rollouts less influence. A larger λ gives rollouts more influence.

Tracing Three Settings

Suppose AlphaGo is evaluating one Go position and you want to understand what λ changes.

Start with λ = 0: Rollouts contribute nothing, so the evaluation is value-network-only.

Move to λ = 0.5: The evaluation uses a midpoint balance: both the value-network evaluation and the rollout evaluation contribute.

Move to λ = 1: The evaluation relies only on rollouts, with no contribution from the value network.

Changing λ changes the source of influence in the final game-state evaluation; it does not create a new evaluation method.

Reading the Endpoint Settings

more rollout influencemore rollout influenceλ = 0Value network onlyλ = 0.5Both methods contributeλ = 1Rollouts only
Which evaluation method contributes when λ equals 0, 0.5, or 1?
SettingInterpretationContributing methods
λ = 0Value-network-only evaluationValue network
λ = 0.5A balanced midpoint between the two evaluation methodsValue network and rollouts
λ = 1Rollout-only evaluationRollouts

The endpoint settings make λ's mixing role easy to interpret.

The endpoints are especially useful because they turn the mixing control into clear comparisons. At λ = 0, rollouts contribute nothing. At λ = 1, the value network contributes nothing. At λ = 0.5, neither method is removed: the final evaluation combines both. The reported best setting was λ = 0.5, showing that the strongest result came from combining the methods rather than choosing either extreme.

Why the Combination Helped

learned viewsimulation-based viewValue networkHigh-performance policy pρComplementaryevaluationCombined with λRolloutsWeaker, faster policy pπ
What is the difference between the value network's learned evaluation and a rollout's simulated game outcome?

AlphaGo assigned different jobs to two imperfectly matched methods. The value network evaluated the high-performance policy pρ, but that policy was too slow for direct live play. Rollouts used the weaker policy pπ, but that policy was faster. In the evaluation process, rollouts could add precision for particular states while the value network supplied its own assessment. λ connected these outputs instead of forcing AlphaGo to discard one of them.

evaluation inputevaluation inputcombined assessmentHigh-performance pρEvaluated by value networkλControls influenceStronger playBest reported setting: λ =0.5Fast pπUsed by rollouts
How does combining a value-network evaluation with rollout evaluation provide a stronger assessment than relying on one method alone?

Mistakes in Interpreting λ

  • Treating λ as a third evaluator

    λ is a mixing control. It determines how the value-network and rollout evaluations contribute to the final evaluation.

    Fix: Think of λ as controlling the balance between two evaluation outputs, not as producing a third output.

  • Reversing the endpoint meanings

    The source defines λ = 0 as value-network-only and λ = 1 as rollout-only.

    Fix: Remember that increasing λ gives rollouts more influence: zero excludes rollouts, while one excludes the value network.

  • Assuming the strongest result must come from one method alone

    Although value-network-only evaluation was stronger than rollout-only evaluation, the best reported play came from λ = 0.5.

    Fix: Compare all three interpretations: the strongest result came from combining the methods at λ = 0.5.

  • Ignoring the different policy roles

    The value network evaluated the high-performance pρ, while rollouts used the weaker but faster pπ.

    Fix: Keep the roles separate: pρ is associated with the value-network evaluation, and pπ is used by rollouts.

Check Your Interpretation

MEDIUM

For each setting, identify which evaluation method contributes: λ = 0, λ = 0.5, and λ = 1. Then explain why the reported best setting, λ = 0.5, does not contradict the result that value-network-only evaluation was stronger than rollout-only evaluation.

Hints
  • Start with the endpoint meanings: λ = 0 excludes rollouts, and λ = 1 excludes the value network.
  • At λ = 0.5, neither evaluation method is excluded.
  • Distinguish the comparison between the two extremes from the result of combining both methods.

What do you think happens?

If λ changes from 0 to 1, which evaluation method gains influence in the final game-state evaluation?

  • The value network
  • Rollouts
  • Neither method
  • A new third method
Reveal answer

Answer: Rollouts

λ controls the balance between the two methods. At λ = 0 the evaluation is value-network-only, while at λ = 1 it is rollout-only.

The Evaluation Strategy

  1. AlphaGo evaluated Go game states with both a value network and rollouts.
  2. The value network evaluated the high-performance policy pρ, while rollouts used the weaker but faster policy pπ.
  3. λ controlled how much the final evaluation relied on each method.
  4. λ = 0 meant value-network-only evaluation, λ = 1 meant rollout-only evaluation, and λ = 0.5 combined both methods.
  5. The best reported play came from λ = 0.5, showing why complementary evaluations could be stronger together than either method alone.

Key Takeaways

  • AlphaGo used a value network and rollouts to evaluate the same Go game state.
  • The value network evaluated the high-performance policy pρ, while rollouts used the weaker but faster policy pπ.
  • λ was a mixing control: λ = 0 selected value-network-only evaluation, and λ = 1 selected rollout-only evaluation.
  • λ = 0.5 combined both evaluation methods, and this was the reported best setting.
  • Combining the methods allowed AlphaGo to benefit from their different strengths instead of relying on only one imperfect evaluation.