Concepts / AlphaGo's Policy and Value Networks

AlphaGo's Policy and Value Networks

AlphaGo evaluated game states with both a value network and rollouts.

  • Programming

Two Views of a Go Position

AlphaGo evaluated a Go game state in two ways rather than relying on a single estimate. One method used a value network. The other used rollouts. The central design question was how much influence each method should have in the final evaluation.

The value network evaluated the high-performance policy pρ. Rollouts used the weaker but faster policy pπ. These methods therefore contributed different strengths: the value network represented a stronger policy, while rollouts could use a faster policy to add precision for particular game states.

evaluateevaluatecontributecontributeGo game stateValue networkpρFinal evaluationRolloutspπ
How do the value network and rollouts receive the same Go game state and contribute to one final estimate?

Following One Position Through the System

A Hypothetical Position

Suppose AlphaGo is evaluating one Go position and must combine the value-network evaluation with the rollout evaluation.

Start with the position: The same Go game state is presented to both evaluation methods.

Use the value network: The value network evaluates the high-performance policy pρ.

Use rollouts: Rollouts evaluate the position using the weaker but faster policy pπ.

Apply λ: λ determines how much the final evaluation relies on the rollout result and how much it relies on the value-network result.

Produce one estimate: The two contributions are combined into the evaluation used for that game state.

The position is judged using both methods, with λ controlling their relative influence.

This example is a mechanism trace, not a numerical calculation. The source describes the roles of the two methods and λ, but it does not provide particular evaluation scores for a specific board position. The important process is that one state receives two evaluations before they are combined.

λ as the Mixing Control

λ is a mixing control, not a third evaluation method. It determines how much the final game-state evaluation relies on the value network and how much it relies on rollouts. Changing λ changes the balance between the two existing evaluations.

usesusesusesusesλ = 0value network onlyValue networkpρλ = 0.5combined evaluationRolloutspπλ = 1rollouts only
As λ changes, how does the influence shift between the value-network evaluation and the rollout evaluation?
λ valueEvaluation usedMeaning
0Value network onlyRollouts contribute nothing
0.5Combined value-network and rollout evaluationBoth methods influence the final evaluation
1Rollout onlyThe evaluation relies only on rollouts

The endpoint meanings come directly from the way λ controls the mixture. The λ = 0.5 row describes the reported best setting as a combined evaluation, not as a claim that the two numerical outputs are necessarily identical.

What the Three Key Settings Mean

What do you think happens?

Which evaluation contributes when λ = 0, and which contributes when λ = 1?

  • λ = 0 uses rollouts only, and λ = 1 uses the value network only
  • λ = 0 uses the value network only, and λ = 1 uses rollouts only
  • Both settings use equal contributions
Reveal answer

Answer: λ = 0 uses the value network only, and λ = 1 uses rollouts only.

At λ = 0, rollouts contribute nothing. At λ = 1, the evaluation relies only on rollouts.

At λ = 0, AlphaGo uses value-network-only evaluation. This removes the rollout contribution. At λ = 1, AlphaGo uses rollout-only evaluation. This removes the value-network contribution. At λ = 0.5, both methods contribute to the final evaluation, representing a balanced combination of the two methods.

Why the Combination Was Stronger

contributescontributessupportsValue networkhigh-performance policy pρλ = 0.5combined evaluationBest playreported settingRolloutsweaker, faster policy pπ
How does combining learned value estimates with rollout outcomes compensate for relying on either method alone?

The value network and rollouts were useful because they were imperfectly matched in a productive way. The value network evaluated the high-performance policy pρ, but that policy was too slow for direct live play. Rollouts used the weaker but faster policy pπ, allowing them to add precision for particular states. Combining their outputs gave AlphaGo more than either isolated evaluation method.

The reported comparison supports this conclusion. AlphaGo using only the value network played better than AlphaGo using only rollouts. However, the best play came from λ = 0.5 rather than from either method alone. The source also reports that the value-network-only version played better than the strongest of the other Go programs mentioned in the evaluation.

Mistakes in Reading λ

  • Reversing the endpoint meanings

    The source defines λ = 0 as value-network-only evaluation and λ = 1 as rollout-only evaluation.

    Fix: Remember that rollouts contribute nothing at λ = 0, while λ = 1 relies only on rollouts.

  • Calling λ a third evaluator

    λ is a mixing control. It determines the relative influence of the value-network and rollout evaluations.

    Fix: Name the two evaluators first, then describe λ as the control that combines their outputs.

  • Assuming one method was always sufficient

    The best reported play came from λ = 0.5, which combined both methods.

    Fix: Distinguish between the stronger endpoint and the best combined setting.

  • Confusing pρ and pπ

    The value network evaluated pρ, while rollouts used the weaker but faster policy pπ.

    Fix: Associate pρ with the value network and pπ with rollouts.

Check Your Understanding

MEDIUM

Explain, in your own words, what changes when λ moves from 0 to 1. Then explain why λ = 0.5 can be stronger than either endpoint even though the value network alone performed better than rollouts alone.

Hints
  • Start by identifying which method is used at each endpoint.
  • Then compare the roles of pρ and pπ.
  • Finally, connect the reported λ = 0.5 result to the complementary strengths of the two methods.

A Complete Explanation

Give a concise explanation of why AlphaGo used both a value network and rollouts.

Identify the value network: It evaluated the high-performance policy pρ.

Identify rollouts: They used the weaker but faster policy pπ.

Identify λ: λ controlled how much each evaluation contributed.

Interpret the result: The reported best play at λ = 0.5 showed that combining the methods was better than relying on either one alone.

AlphaGo gained strength by combining a high-performance value-network evaluation with a faster rollout evaluation instead of selecting only one.

Key Takeaways

  1. AlphaGo evaluated a Go state with both a value network and rollouts.
  2. The value network evaluated the high-performance policy pρ, while rollouts used the weaker but faster policy pπ.
  3. λ controlled the mixture: λ = 0 meant value-network-only evaluation, and λ = 1 meant rollout-only evaluation.
  4. λ = 0.5 combined both evaluation methods and produced the best reported play.
  5. The combination worked because the two imperfect methods contributed complementary strengths.

Key Takeaways

  • AlphaGo used a value network and rollouts to evaluate the same Go game state.
  • The value network evaluated pρ, the high-performance policy; rollouts used pπ, the weaker but faster policy.
  • λ controlled the balance between the two evaluations.
  • λ = 0 used only the value network, λ = 1 used only rollouts, and λ = 0.5 combined them.
  • The reported best play at λ = 0.5 demonstrated the strength of combining complementary evaluation methods.