Concepts / Policy Evaluation in Tic-Tac-Toe

Policy Evaluation in Tic-Tac-Toe

Evolutionary methods evaluate a fixed policy through many game outcomes.

  • Programming

One Policy Under Test

Suppose you have a policy for playing Tic-Tac-Toe. The policy is the behavior being evaluated: it determines how play proceeds when the policy is used. The central question is not whether one move looks good in isolation. The question is how well the policy performs overall. Evolutionary methods answer that question by holding one policy fixed, using it in many games, and judging it from the outcomes of those games.

An evolutionary evaluation follows this broad path: fixed policy, many games, final outcomes, win frequency, later policy selection.

is tested inproduceare summarized asFixed policyMany gamesGame outcomeswins and other finalresultsWin frequencyperformance estimate
How do repeated game outcomes flow into an estimate of how well one fixed policy performs?

Repeated Games

The policy remains unchanged while it is being tested. It plays many games against an opponent. The opponent can be an actual opponent, or the games can be simulated with a model of the opponent. Each game ends with an outcome, and the policy is judged from the collection of those final outcomes rather than from a separate evaluation of every board position encountered during play.

Evaluating One Fixed Policy

Imagine testing one Tic-Tac-Toe policy without changing it during the test.

Hold the policy fixed: Use the same policy throughout the evaluation. No policy change occurs in the middle of this test.

Play many games: Let the fixed policy play many games against an opponent, either through actual play or through simulation with a model of the opponent.

Collect final outcomes: Record the outcomes of the completed games. The evaluation is based on these game outcomes.

Estimate performance: Use the frequency of wins as an estimate of the probability of winning for the policy in this evaluation.

The policy receives a broad, game-level evaluation rather than a separate evaluation of each position visited during play.

Win Frequency

Win frequency is the proportion of tested games that the policy wins. It estimates the probability of winning and gives a basis for choosing what happens next. A policy with a higher observed win frequency is treated as having stronger performance in that evaluation, so its result can guide the selection of a later policy.

Comparing Candidate Policies

Three candidate policies are evaluated over repeated games. Policy A wins 6 of 10 games, Policy B wins 3 of 10 games, and Policy C wins 8 of 10 games.

Summarize each evaluation: Each policy is represented by the frequency of its wins across its tested games.

Compare the frequencies: Policy C has the highest win frequency in this generated example, followed by Policy A and then Policy B.

Use the comparison: The higher win frequency of Policy C provides the strongest outcome-based evidence among these three candidates for selecting or generating a later policy.

Policy C is the candidate selected by the win-frequency comparison in this illustrative example.

reportsreportsreportsguidesPolicy Awin frequencyFrequency comparisonLater policyselected or generatedPolicy Bwin frequencyPolicy Cwin frequency
How do win frequencies from several candidate policies determine which policy is selected or generated next?

A policy change comes after the games used for evaluation. This separates testing the current policy from selecting or generating a later policy.

State-by-State Evaluation

A value function method follows a different trace. Instead of waiting for a collection of complete-game outcomes to represent the evaluation, it allows individual states to be evaluated during play. In Tic-Tac-Toe, a state is an individual board position encountered as the game progresses. The method therefore uses information that appears during the game.

play progressesplay progresseseventually reachesBoard state 1value evaluationBoard state 2value evaluationBoard state 3value evaluationGame outcome
How does a value function assign different evaluations to board states as the game progresses?

The important contrast is where the evaluation information enters. Evolutionary evaluation assigns credit from final outcomes across games. Value function evaluation uses evaluations of individual states encountered during play. Thus, the value-function trace includes intermediate positions instead of representing the policy only through the final results of a collection of games.

Two Kinds of Information

Evolutionary methodsValue function methods
Evaluate a fixed policy through many game outcomesEvaluate individual states during play
Use final outcomes from completed gamesUse information that appears during the game
Summarize performance with win frequencyRepresent evaluation at the state level
Policy change comes after many gamesEvaluation can occur as individual states are visited
summarizesassignsEvolutionary methodaggregate game outcomesWin frequencypolicy-level estimateValue functionmethodindividual game statesState evaluationduring-play information
What information does each method use: aggregate outcomes across complete games or evaluations of individual states during play?

Searching for Better Policies

Both approaches can be understood as ways to search through possible policies, but they organize the search around different evidence. An evolutionary method evaluates a policy as a whole, uses the frequency of wins to guide later policy selection, and changes policy after many games. A value function method brings information from individual states into the evaluation during play. That state-level information can guide the search for policy behavior using what is learned about positions encountered in games.

is tested inproduceguidesencountersreceiveinformsPolicywhole-policy evaluationMany gamesfixed policyWin frequencypolicy comparisonLater policyPolicystate-informed behaviorIndividual statesduring playState evaluationsplay informationLater policy
How does each approach move from one policy to another while searching for better play?

The shared goal is policy improvement through search. The difference is the evidence used to navigate that search: aggregate game performance for evolutionary methods and during-play state evaluations for value function methods.

Common Mistakes

  • Treating evolutionary evaluation as an evaluation of every board position.

    The evolutionary method judges the fixed policy from final outcomes across many games and ignores what happens during games when assigning credit.

    Fix: Describe the evaluation path as policy to games to final outcomes to win frequency.

  • Changing the policy while measuring that same policy.

    The method holds the tested policy fixed while the games are played.

    Fix: Complete the evaluation of the fixed policy before using the result to guide a later policy.

  • Confusing win frequency with a value assigned to one state.

    Win frequency is a policy-level estimate based on game outcomes, while value function methods evaluate individual states during play.

    Fix: Use win frequency when discussing the evolutionary evaluation of a policy and state evaluation when discussing value function methods.

  • Assuming both methods use the same timing of information.

    Value function methods use information that appears during play by evaluating individual states.

    Fix: Compare final-outcome information with during-play state information.

Practice

MEDIUM

A Tic-Tac-Toe policy is held fixed while it plays many simulated games. After the games finish, the evaluator summarizes how often the policy won. Identify the evaluation method, name the information used to judge the policy, and explain how that result can affect the next policy.

Hints
  • Look for whether the policy changes during testing.
  • Identify whether the evidence comes from complete-game outcomes or individual states.
  • Connect the frequency of wins to later policy selection.

What do you think happens?

A method evaluates board states as they are encountered during a Tic-Tac-Toe game instead of waiting to represent the policy only through final results from many games. Which approach does this describe?

  • An evolutionary method
  • A value function method
  • Neither approach
Reveal answer

Answer: A value function method

Value function methods evaluate individual states during play and therefore use information that appears during the game.

Summary

  1. Evolutionary methods evaluate one fixed Tic-Tac-Toe policy through many game outcomes. The frequency of wins estimates the policy's probability of winning and guides the selection or generation of a later policy. Value function methods evaluate individual states during play, so they use information available before a collection of complete games has been reduced to final outcomes. Both approaches can participate in a search for better policies, but they use different evaluation traces: evolutionary methods use aggregate game outcomes, while value function methods use state-level information encountered during play.

Key Takeaways

  • A fixed policy is evaluated by playing many games and collecting their final outcomes.
  • Win frequency estimates the probability of winning and guides later policy selection.
  • Value function methods evaluate individual board states during play.
  • Evolutionary methods use aggregate outcomes, while value function methods use information from states encountered during games.
  • Both approaches can search for better policies, but they organize that search around different kinds of evidence.