Concepts / Blackjack Policy Evaluation

Blackjack Policy Evaluation

Monte Carlo ES turns simulated Blackjack episodes into a search for a strong policy.

  • Programming

From Games to Policy Improvement

A Blackjack strategy can be judged by playing many complete games, but Monte Carlo ES does more than evaluate one fixed strategy. It uses simulated games to search for better action choices. The central loop is: arrange simulated starts so that different Blackjack situations can occur, observe the results of complete episodes, use those results to update action values, and improve the policy using the updated values.

begincompleteupdateimproverepeatExploring startsBlackjack situations andactionsSimulated episodeComplete gameObserved resultsEpisode returnsAction valuesChoices in situationsImproved policyAction choice
How does a simulated Blackjack episode produce returns that update action values and improve the policy?

The important distinction is between evaluation and improvement. Evaluation estimates how well choices perform. Monte Carlo ES uses those estimates to search for a stronger policy rather than merely reporting the performance of a policy that was fixed in advance.

Coverage Through Exploring Starts

Exploring starts address a coverage problem. If every simulated game began in only a narrow range of situations, some player sums, dealer cards, or usable-ace conditions might receive little or no attention. Because the games are simulated, their starts can be arranged directly. The dealer's cards, the player's sum, and whether the player has a usable ace are each selected at random with equal probability. This makes all of these possibilities eligible to appear as starting situations.

selectselectselectforms situationforms situationforms situationRandom startEqual probabilityDealer cardsSelected possibilityPlayer's sumSelected possibilityUsable aceSelected conditionStarting actionAction choice
How do randomly selected starting situations ensure that different Blackjack possibilities can be explored?

Why the start matters

Consider why a simulation would need more than one narrow kind of Blackjack beginning.

Identify the situation dimensions: A Blackjack situation is organized by the player's sum, the dealer's showing card, and whether the player has a usable ace.

Identify the coverage risk: If starts were restricted, some combinations of these dimensions might receive little or no attention.

Apply exploring starts: The simulated episode can begin by selecting the dealer's cards, the player's sum, and the usable-ace condition at random with equal probability.

Connect coverage to improvement: Different situations can then contribute observed results to action values, giving the policy a basis for choosing actions across the described possibilities.

Exploring starts turn the beginning of each simulated game into a coverage mechanism rather than allowing the simulations to focus only on a narrow set of situations.

Values Behind the Policy

The policy is organized around the information available in a Blackjack situation: the player's sum, the dealer's showing card, and whether the player has a usable ace. For each such situation, the policy identifies an action such as hit or stick.

Action values describe choices made in particular situations. They associate a situation with an action, such as hitting or sticking, and represent the value used when comparing those choices. The state-value function is obtained from the action values. In other words, the value of a situation is derived from the action-value information available for the actions under the policy, rather than being an unrelated estimate.

considerconsidercandidatecandidatederiveBlackjack stateSum, dealer card, aceconditionHitAction valueStickAction valuePolicy actionSelected choiceState valueDerived from action values
How is the value of a Blackjack state derived from the action values for hitting and sticking under the learned policy?

Tracing One Improvement Cycle

A conceptual Monte Carlo ES cycle

Trace what happens when one simulated Blackjack episode begins from an exploring start.

Begin with the starting policy: The example begins with a policy that sticks on 20 or 21. This is only the starting policy, not the final reported result.

Arrange the episode start: The dealer's cards, the player's sum, and the usable-ace condition are selected at random with equal probability so that the situation is eligible for exploration.

Play the complete episode: The simulated game produces an observed result that can be used to evaluate the action choices encountered.

Update action values: The observed results contribute to action-value estimates for situation-action choices, including choices such as hit or stick.

Improve the policy: The updated action-value information is used to search for stronger action choices. Repeating this process produces the learned policy.

Monte Carlo ES converts complete simulated games into policy improvement by combining broad starts, observed results, action-value updates, and revised action choices.

compared withcompared withhas differencedoes not showMonte Carlo ESpolicyOptimal for described gameThorp's strategyBasic strategySimilar policyBroad agreementLeftmost notchUsable-ace states
Where does the Monte Carlo ES policy agree with Thorp's basic strategy, and where does the source report a difference?
Policy comparisonWhat the source reports
Overall relationshipThe Monte Carlo ES policy is similar to Thorp's basic strategy.
Reported differenceA leftmost notch in the policy for states with a usable ace is absent from Thorp's strategy.
Explanation of the differenceThe source does not establish why this difference occurs.
Status of the learned policyThe policy is reported as optimal for the particular version of Blackjack described.

Check Your Understanding

MEDIUM

Explain why Monte Carlo ES uses exploring starts instead of relying only on the ordinary beginnings of simulated games. Then describe the sequence from an episode's observed result to an improved policy.

Hints
  • Mention the player's sum, the dealer's cards, and the usable-ace condition.
  • Distinguish action values from the state-value function.
  • Remember that the starting stick-on-20-or-21 policy is not the final learned policy.
  • Treating Monte Carlo ES as only an evaluation method.

    The method uses simulated games to search for a better policy, not merely to evaluate one fixed strategy.

    Fix: Track both stages: observed episode results update action values, and those values support policy improvement.

  • Thinking that exploring starts mean randomizing only the first action.

    The source describes random selection of the dealer's cards, the player's sum, and the usable-ace condition.

    Fix: Treat exploring starts as a coverage arrangement for Blackjack situations, with the listed elements selected at random with equal probability.

  • Confusing an action value with a state value.

    Action values describe state-action choices, while the state-value function is obtained from those action values.

    Fix: Keep the situation and the chosen action conceptually separate, then derive the state value from the relevant action-value information.

  • Assuming that the starting stick-on-20-or-21 policy is the learned result.

    The source describes it as the starting policy and reports that Monte Carlo ES later produces an optimal policy for the described game.

    Fix: Label the stick-on-20-or-21 rule as the initial policy and the Monte Carlo ES result as the learned policy.

  • Claiming a reason for the difference from Thorp's strategy.

    The source identifies the discrepancy but does not establish its reason.

    Fix: Report the difference without adding an unsupported cause.

Key Takeaways

  1. Monte Carlo ES uses complete simulated Blackjack games to search for a stronger policy, not merely to evaluate one fixed strategy.
  2. Exploring starts provide coverage by randomly selecting the dealer's cards, the player's sum, and the usable-ace condition with equal probability.
  3. The Blackjack example begins with a stick-on-20-or-21 policy; the supplied source does not state numerical initial action-value estimates.
  4. Action values describe state-action choices, and the state-value function is derived from those action values.
  5. The learned policy is reported as optimal for the described game and is similar to Thorp's basic strategy, with the source noting one difference involving a leftmost notch for usable-ace states.

Key Takeaways

  • Monte Carlo ES turns simulated Blackjack episodes into policy improvement.
  • Exploring starts ensure that different combinations of player sum, dealer cards, and usable-ace condition can be encountered.
  • The example starts with a stick-on-20-or-21 policy; the supplied source does not specify numerical initial action-value estimates.
  • State values are obtained from action values for the available choices.
  • The learned policy is similar to Thorp's basic strategy, with one reported difference for usable-ace states.