Policy Evaluation in Tic-Tac-Toe
Evolutionary methods evaluate a fixed policy through many game outcomes.
One Policy Under Test
Suppose you have a policy for playing Tic-Tac-Toe. The policy is the behavior being evaluated: it determines how play proceeds when the policy is used. The central question is not whether one move looks good in isolation. The question is how well the policy performs overall. Evolutionary methods answer that question by holding one policy fixed, using it in many games, and judging it from the outcomes of those games.
An evolutionary evaluation follows this broad path: fixed policy, many games, final outcomes, win frequency, later policy selection.
Repeated Games
The policy remains unchanged while it is being tested. It plays many games against an opponent. The opponent can be an actual opponent, or the games can be simulated with a model of the opponent. Each game ends with an outcome, and the policy is judged from the collection of those final outcomes rather than from a separate evaluation of every board position encountered during play.
Evaluating One Fixed Policy
Imagine testing one Tic-Tac-Toe policy without changing it during the test.
Hold the policy fixed: Use the same policy throughout the evaluation. No policy change occurs in the middle of this test.
Play many games: Let the fixed policy play many games against an opponent, either through actual play or through simulation with a model of the opponent.
Collect final outcomes: Record the outcomes of the completed games. The evaluation is based on these game outcomes.
Estimate performance: Use the frequency of wins as an estimate of the probability of winning for the policy in this evaluation.
The policy receives a broad, game-level evaluation rather than a separate evaluation of each position visited during play.
Win Frequency
Win frequency is the proportion of tested games that the policy wins. It estimates the probability of winning and gives a basis for choosing what happens next. A policy with a higher observed win frequency is treated as having stronger performance in that evaluation, so its result can guide the selection of a later policy.
Comparing Candidate Policies
Three candidate policies are evaluated over repeated games. Policy A wins 6 of 10 games, Policy B wins 3 of 10 games, and Policy C wins 8 of 10 games.
Summarize each evaluation: Each policy is represented by the frequency of its wins across its tested games.
Compare the frequencies: Policy C has the highest win frequency in this generated example, followed by Policy A and then Policy B.
Use the comparison: The higher win frequency of Policy C provides the strongest outcome-based evidence among these three candidates for selecting or generating a later policy.
Policy C is the candidate selected by the win-frequency comparison in this illustrative example.
A policy change comes after the games used for evaluation. This separates testing the current policy from selecting or generating a later policy.
State-by-State Evaluation
A value function method follows a different trace. Instead of waiting for a collection of complete-game outcomes to represent the evaluation, it allows individual states to be evaluated during play. In Tic-Tac-Toe, a state is an individual board position encountered as the game progresses. The method therefore uses information that appears during the game.
The important contrast is where the evaluation information enters. Evolutionary evaluation assigns credit from final outcomes across games. Value function evaluation uses evaluations of individual states encountered during play. Thus, the value-function trace includes intermediate positions instead of representing the policy only through the final results of a collection of games.
Two Kinds of Information
| Evolutionary methods | Value function methods |
|---|---|
| Evaluate a fixed policy through many game outcomes | Evaluate individual states during play |
| Use final outcomes from completed games | Use information that appears during the game |
| Summarize performance with win frequency | Represent evaluation at the state level |
| Policy change comes after many games | Evaluation can occur as individual states are visited |
Searching for Better Policies
Both approaches can be understood as ways to search through possible policies, but they organize the search around different evidence. An evolutionary method evaluates a policy as a whole, uses the frequency of wins to guide later policy selection, and changes policy after many games. A value function method brings information from individual states into the evaluation during play. That state-level information can guide the search for policy behavior using what is learned about positions encountered in games.
The shared goal is policy improvement through search. The difference is the evidence used to navigate that search: aggregate game performance for evolutionary methods and during-play state evaluations for value function methods.
Common Mistakes
Treating evolutionary evaluation as an evaluation of every board position.
The evolutionary method judges the fixed policy from final outcomes across many games and ignores what happens during games when assigning credit.
Fix:
Describe the evaluation path as policy to games to final outcomes to win frequency.Changing the policy while measuring that same policy.
The method holds the tested policy fixed while the games are played.
Fix:
Complete the evaluation of the fixed policy before using the result to guide a later policy.Confusing win frequency with a value assigned to one state.
Win frequency is a policy-level estimate based on game outcomes, while value function methods evaluate individual states during play.
Fix:
Use win frequency when discussing the evolutionary evaluation of a policy and state evaluation when discussing value function methods.Assuming both methods use the same timing of information.
Value function methods use information that appears during play by evaluating individual states.
Fix:
Compare final-outcome information with during-play state information.
Practice
A Tic-Tac-Toe policy is held fixed while it plays many simulated games. After the games finish, the evaluator summarizes how often the policy won. Identify the evaluation method, name the information used to judge the policy, and explain how that result can affect the next policy.
Hints
- Look for whether the policy changes during testing.
- Identify whether the evidence comes from complete-game outcomes or individual states.
- Connect the frequency of wins to later policy selection.
What do you think happens?
A method evaluates board states as they are encountered during a Tic-Tac-Toe game instead of waiting to represent the policy only through final results from many games. Which approach does this describe?
Reveal answer
Answer: A value function method
Value function methods evaluate individual states during play and therefore use information that appears during the game.
Summary
- Evolutionary methods evaluate one fixed Tic-Tac-Toe policy through many game outcomes. The frequency of wins estimates the policy's probability of winning and guides the selection or generation of a later policy. Value function methods evaluate individual states during play, so they use information available before a collection of complete games has been reduced to final outcomes. Both approaches can participate in a search for better policies, but they use different evaluation traces: evolutionary methods use aggregate game outcomes, while value function methods use state-level information encountered during play.
Key Takeaways
- A fixed policy is evaluated by playing many games and collecting their final outcomes.
- Win frequency estimates the probability of winning and guides later policy selection.
- Value function methods evaluate individual board states during play.
- Evolutionary methods use aggregate outcomes, while value function methods use information from states encountered during games.
- Both approaches can search for better policies, but they organize that search around different kinds of evidence.