Concepts / AlphaGo's Value Network

AlphaGo's Value Network

AlphaGo evaluated game states with both a value network and rollouts.

  • Programming

Two Views of a Go Position

When AlphaGo considered a Go game state, it did not rely on only one evaluation method. It used both a value network and rollouts. The value network and the rollouts gave AlphaGo two complementary ways to judge the same position.

evaluateevaluatecontributecontributeGo game stateValue networkpρCombined evaluationRolloutspπ
How do the value network and rollouts independently evaluate the same Go game state before their results are combined?

The value network evaluated the high-performance policy pρ. Rollouts used the weaker but faster policy pπ. These were not two names for the same operation: they were separate evaluation paths. AlphaGo's design gave each method a role, then used λ to decide how much influence each output would have on the final evaluation.

Following One State Through the System

A Candidate Position

Imagine that AlphaGo is evaluating one candidate Go game state.

First path: The value network evaluates the state using the high-performance policy pρ.

Second path: Rollouts evaluate the same state using the weaker but faster policy pπ.

Mixing step: The two evaluation outputs are mixed. λ determines how much the rollout result contributes and how much the value-network result contributes.

Final judgment: AlphaGo uses the mixed evaluation when judging the candidate state.

The state receives a combined evaluation rather than a judgment from only one method.

The important point is that the state is evaluated twice through different methods before the results are combined. The value network supplies one assessment, while rollouts supply another. λ does not perform a third evaluation; it is the control that mixes the two existing evaluations.

Changing the Mixing Control

Changing λ changes the relative influence of the two evaluation methods. At one endpoint, rollouts contribute nothing and the value network supplies the evaluation. At the other endpoint, the evaluation relies only on rollouts. Values between these endpoints blend the two methods.

increase rollout influenceincrease rollout influenceλ = 0value network onlyλ = 0.5combined evaluationλ = 1rollouts only
How does changing λ alter the relative contribution of the value network and rollouts to the final game-state evaluation?
λ settingValue-network contributionRollout contributionMeaning
λ = 0All of the evaluationNoneValue-network-only evaluation
λ = 0.5Part of the evaluationPart of the evaluationA combined evaluation using both methods
λ = 1NoneAll of the evaluationRollout-only evaluation

Interpretation of the key λ settings

The middle setting λ = 0.5 is especially important in the reported results. It represents a combined evaluation, not a choice between the methods. AlphaGo's best play came from this combined setting rather than from either endpoint.

Why the Combination Worked

has limitationaddscontributescontributessupportsValue networkhigh-performance policy pρToo slow for directlive playCombined evaluationλ = 0.5Stronger playRolloutsweaker but faster policy pπPrecision forparticular states
What complementary strengths do the value network and rollouts contribute, and how does combining them produce a stronger evaluation?

The two methods were imperfectly matched in a useful way. The value network evaluated the high-performance policy pρ, but that policy was too slow for direct live play. Rollouts used the weaker but faster policy pπ and added precision for particular states. Combining their outputs allowed AlphaGo to use information from both methods instead of accepting the limitations of either method alone.

The reported comparison supports this interpretation. AlphaGo using only the value network played better than AlphaGo using only rollouts. It also played better than the strongest of the other Go programs mentioned in the evaluation. However, the best play came from λ = 0.5. The lesson is not that one method was useless; it is that their combination produced the strongest reported evaluation.

Mistakes in Reading λ

  • Treating λ as a third evaluation method

    λ controls how the value-network and rollout evaluations are mixed; it does not evaluate the position on its own.

    Fix: Read λ as a mixing control between two existing evaluation outputs.

  • Reversing the endpoint meanings

    The source defines λ = 0 as value-network-only evaluation and λ = 1 as rollout-only evaluation.

    Fix: Remember that increasing λ moves the evaluation toward rollouts.

  • Assuming the strongest single method must be the best overall setting

    The reported best play came from λ = 0.5, even though value-network-only evaluation outperformed rollout-only evaluation.

    Fix: Distinguish between comparing each method alone and combining both methods.

  • Assuming rollouts used the same policy as the value network

    The value network evaluated the high-performance policy pρ, while rollouts used the weaker but faster policy pπ.

    Fix: Keep the policy roles separate: pρ belongs to the value-network evaluation, and pπ belongs to rollouts.

Check Your Interpretation

MEDIUM

Explain what happens to the evaluation when λ changes from 0 to 0.5 and then to 1. In your answer, identify which method contributes at each endpoint and explain why λ = 0.5 is not the same as selecting only the better standalone method.

Hints
  • Start with the definitions of λ = 0 and λ = 1.
  • Describe λ = 0.5 as a combination rather than an endpoint.
  • Use the reported playing-strength comparison to explain why the mixture matters.

What do you think happens?

If λ is set to 1, which evaluation method remains?

  • The value network only
  • Rollouts only
  • An equal mixture of both methods
Reveal answer

Answer: Rollouts only

λ = 1 means that the evaluation relies only on rollouts. λ = 0 is the value-network-only endpoint, while λ = 0.5 combines both methods.

Key Takeaways

  1. AlphaGo evaluated a Go game state with both a value network and rollouts.
  2. The value network evaluated the high-performance policy pρ, while rollouts used the weaker but faster policy pπ.
  3. λ controlled the mixture: λ = 0 meant value-network-only evaluation, and λ = 1 meant rollout-only evaluation.
  4. λ = 0.5 produced a combined evaluation and was the reported setting for AlphaGo's best play.
  5. Combining the methods was effective because their different capabilities complemented one another.

Key Takeaways

  • AlphaGo used a value network and rollouts as separate ways to evaluate the same Go game state.
  • The value network used the high-performance policy pρ, while rollouts used the weaker but faster policy pπ.
  • λ controlled the balance between the two evaluations: 0 means value-network-only, 1 means rollout-only, and 0.5 means a combined evaluation.
  • The strongest reported play came from combining the methods at λ = 0.5 rather than relying on either method alone.