AlphaGo's Rollout Policy
AlphaGo evaluated game states with both a value network and rollouts.
Two Views of One Go Position
When AlphaGo evaluated a Go game state, it did not rely on just one kind of judgment. It used a value network and rollouts as complementary evaluation methods. The value network evaluated the high-performance policy pρ, while rollouts used the weaker but faster policy pπ. The central design question was how much influence each evaluation should have.
The two paths answer the same broad question: how should this Go position be evaluated? They do so through different mechanisms. The value network provides the learned evaluation associated with the high-performance policy pρ. Rollouts provide an evaluation based on the weaker but faster policy pπ. AlphaGo's rollout policy was therefore not a replacement for the value network. It was a second source of information about the same position.
Following λ Through the Evaluation
AlphaGo used λ as a mixing control. It determined how much the final game-state evaluation relied on the value network and how much it relied on rollouts. λ was not a third evaluation method and did not produce an independent judgment about the Go position. Its role was to control the balance between the two existing judgments.
To interpret λ, ask one question: how much of the final evaluation is being assigned to rollouts rather than to the value network? A smaller λ gives rollouts less influence. A larger λ gives rollouts more influence.
Tracing Three Settings
Suppose AlphaGo is evaluating one Go position and you want to understand what λ changes.
Start with λ = 0: Rollouts contribute nothing, so the evaluation is value-network-only.
Move to λ = 0.5: The evaluation uses a midpoint balance: both the value-network evaluation and the rollout evaluation contribute.
Move to λ = 1: The evaluation relies only on rollouts, with no contribution from the value network.
Changing λ changes the source of influence in the final game-state evaluation; it does not create a new evaluation method.
Reading the Endpoint Settings
| Setting | Interpretation | Contributing methods |
|---|---|---|
| λ = 0 | Value-network-only evaluation | Value network |
| λ = 0.5 | A balanced midpoint between the two evaluation methods | Value network and rollouts |
| λ = 1 | Rollout-only evaluation | Rollouts |
The endpoint settings make λ's mixing role easy to interpret.
The endpoints are especially useful because they turn the mixing control into clear comparisons. At λ = 0, rollouts contribute nothing. At λ = 1, the value network contributes nothing. At λ = 0.5, neither method is removed: the final evaluation combines both. The reported best setting was λ = 0.5, showing that the strongest result came from combining the methods rather than choosing either extreme.
Why the Combination Helped
AlphaGo assigned different jobs to two imperfectly matched methods. The value network evaluated the high-performance policy pρ, but that policy was too slow for direct live play. Rollouts used the weaker policy pπ, but that policy was faster. In the evaluation process, rollouts could add precision for particular states while the value network supplied its own assessment. λ connected these outputs instead of forcing AlphaGo to discard one of them.
Mistakes in Interpreting λ
Treating λ as a third evaluator
λ is a mixing control. It determines how the value-network and rollout evaluations contribute to the final evaluation.
Fix:
Think of λ as controlling the balance between two evaluation outputs, not as producing a third output.Reversing the endpoint meanings
The source defines λ = 0 as value-network-only and λ = 1 as rollout-only.
Fix:
Remember that increasing λ gives rollouts more influence: zero excludes rollouts, while one excludes the value network.Assuming the strongest result must come from one method alone
Although value-network-only evaluation was stronger than rollout-only evaluation, the best reported play came from λ = 0.5.
Fix:
Compare all three interpretations: the strongest result came from combining the methods at λ = 0.5.Ignoring the different policy roles
The value network evaluated the high-performance pρ, while rollouts used the weaker but faster pπ.
Fix:
Keep the roles separate: pρ is associated with the value-network evaluation, and pπ is used by rollouts.
Check Your Interpretation
For each setting, identify which evaluation method contributes: λ = 0, λ = 0.5, and λ = 1. Then explain why the reported best setting, λ = 0.5, does not contradict the result that value-network-only evaluation was stronger than rollout-only evaluation.
Hints
- Start with the endpoint meanings: λ = 0 excludes rollouts, and λ = 1 excludes the value network.
- At λ = 0.5, neither evaluation method is excluded.
- Distinguish the comparison between the two extremes from the result of combining both methods.
What do you think happens?
If λ changes from 0 to 1, which evaluation method gains influence in the final game-state evaluation?
Reveal answer
Answer: Rollouts
λ controls the balance between the two methods. At λ = 0 the evaluation is value-network-only, while at λ = 1 it is rollout-only.
The Evaluation Strategy
- AlphaGo evaluated Go game states with both a value network and rollouts.
- The value network evaluated the high-performance policy pρ, while rollouts used the weaker but faster policy pπ.
- λ controlled how much the final evaluation relied on each method.
- λ = 0 meant value-network-only evaluation, λ = 1 meant rollout-only evaluation, and λ = 0.5 combined both methods.
- The best reported play came from λ = 0.5, showing why complementary evaluations could be stronger together than either method alone.
Key Takeaways
- AlphaGo used a value network and rollouts to evaluate the same Go game state.
- The value network evaluated the high-performance policy pρ, while rollouts used the weaker but faster policy pπ.
- λ was a mixing control: λ = 0 selected value-network-only evaluation, and λ = 1 selected rollout-only evaluation.
- λ = 0.5 combined both evaluation methods, and this was the reported best setting.
- Combining the methods allowed AlphaGo to benefit from their different strengths instead of relying on only one imperfect evaluation.