AlphaGo’s Supervised-Learning Policy Network
The reinforcement-learning policy network began with the final weights of the supervised-learning policy network.
From Imitation to Competition
AlphaGo’s reinforcement-learning stage did not begin with an unrelated network. It reused the same architecture as the supervised-learning policy network and began with that network’s final learned weights. Reinforcement learning then improved this existing policy by making it compete in simulated games and using the game outcomes as reward signals.
One Simulated Game
The training loop can be followed through one simulated game. The current policy plays directly against an opponent selected from policies produced by earlier iterations. Search does not decide the moves in this competition. After the game reaches an outcome, the current policy receives a numerical reward based on that outcome.
The opponent is not always the current policy. Earlier policies are randomly selected so the current policy is not trained only against itself. The source describes this collection of opponents as a way to prevent overfitting to the current policy by broadening the learning experience.
Rewards from the Current Policy’s View
The reward is assigned from the current policy’s point of view. A win gives the current policy +1, a loss gives it -1, and any other outcome receives 0.
Reading Rewards Correctly
A current policy plays a simulated game against a randomly selected earlier policy. Determine the reward from the current policy’s perspective for three possible outcomes.
Current policy wins: The reward is +1 because the reference point is always the current policy.
Current policy loses: The reward is -1 because the current policy did not win.
The outcome is neither a win nor a loss: The reward is 0 because the source assigns 0 otherwise.
The reward mapping is win to +1, loss to -1, and otherwise to 0.
Improvement versus Evaluation
AlphaGo’s reinforcement-learning stage contains two different purposes. Policy gradient reinforcement learning improves the policy network. It uses simulated games and their reward signals to produce the reinforcement-learning policy, described in the source as p rho. Monte-Carlo policy evaluation comes afterward and estimates the value function of the improved policy.
| Stage | Main question | Result |
|---|---|---|
| Policy gradient reinforcement learning | How can the policy be improved using simulated games and rewards? | An improved reinforcement-learning policy |
| Monte-Carlo policy evaluation | How good is the resulting policy? | An estimated value function |
The two stages are related but perform different jobs.
Value Estimates for APV-MCTS
After policy improvement, Monte-Carlo policy evaluation estimates the value function associated with the resulting policy. That value function is the output needed by APV-MCTS to evaluate positions. The important sequence is therefore: improve the policy through rewarded simulated games, evaluate the resulting policy, and use the estimated value function in APV-MCTS.
Common Reasoning Errors
Treating reinforcement learning as if it started from an unrelated network
The reinforcement-learning policy network began with the final weights of the supervised-learning policy network and used the same architecture.
Fix:
Think of supervised learning as providing the starting policy, which reinforcement learning then improves.Assuming the current policy always plays against itself
Training used randomly selected policies from earlier iterations as opponents.
Fix:
Track the opponent source: it is selected from earlier policy iterations.Assigning reward from the opponent’s perspective
Rewards are defined for the current policy: +1 for its win and -1 for its loss.
Fix:
Name the current policy first, then evaluate the outcome from that policy’s point of view.Confusing policy improvement with policy evaluation
Policy gradient reinforcement learning improves the policy, while Monte-Carlo policy evaluation estimates the value function of the resulting policy.
Fix:
Associate improvement with policy gradient reinforcement learning and value estimation with Monte-Carlo policy evaluation.
Check Your Understanding
Explain the complete training sequence in your own words. Begin with the supervised-learning policy network’s final weights, describe how a current policy is matched with an earlier policy in simulated games, assign rewards for the possible outcomes, and finish by explaining how Monte-Carlo policy evaluation provides the value function needed by APV-MCTS.
Hints
- Keep the current policy as the reference point when assigning rewards.
- Mention why a collection of earlier opponents is used instead of only the current policy.
- Separate the stage that improves the policy from the stage that estimates its value function.
What do you think happens?
A simulated game ends with a loss for the current policy. What reward should the current policy receive?
Reveal answer
Answer: -1
The reward is defined from the current policy’s perspective: a win gives +1, a loss gives -1, and 0 is assigned otherwise.
Key Takeaways
- The reinforcement-learning policy network started with the final weights and the same architecture of the supervised-learning policy network.
- Simulated games matched the current policy against randomly selected policies from earlier iterations, helping prevent overfitting to the current policy.
- From the current policy’s perspective, a win receives +1, a loss receives -1, and any other outcome receives 0.
- Policy gradient reinforcement learning improved the policy, while Monte-Carlo policy evaluation estimated the value function of the improved policy.
- The evaluated policy’s value function supplied the values needed by APV-MCTS to evaluate positions.
Key Takeaways
- Reinforcement learning began from the supervised-learning policy network’s final weights rather than from an unrelated network.
- The current policy played simulated games against randomly selected policies from earlier iterations to broaden training and prevent overfitting to the current policy.
- Rewards were +1 for a current-policy win, -1 for a current-policy loss, and 0 otherwise.
- Policy improvement and Monte-Carlo policy evaluation were separate stages with separate purposes.
- Monte-Carlo evaluation estimated the value function later used by APV-MCTS.