Blackjack Policy Evaluation
Monte Carlo ES turns simulated Blackjack episodes into a search for a strong policy.
From Games to Policy Improvement
A Blackjack strategy can be judged by playing many complete games, but Monte Carlo ES does more than evaluate one fixed strategy. It uses simulated games to search for better action choices. The central loop is: arrange simulated starts so that different Blackjack situations can occur, observe the results of complete episodes, use those results to update action values, and improve the policy using the updated values.
The important distinction is between evaluation and improvement. Evaluation estimates how well choices perform. Monte Carlo ES uses those estimates to search for a stronger policy rather than merely reporting the performance of a policy that was fixed in advance.
Coverage Through Exploring Starts
Exploring starts address a coverage problem. If every simulated game began in only a narrow range of situations, some player sums, dealer cards, or usable-ace conditions might receive little or no attention. Because the games are simulated, their starts can be arranged directly. The dealer's cards, the player's sum, and whether the player has a usable ace are each selected at random with equal probability. This makes all of these possibilities eligible to appear as starting situations.
Why the start matters
Consider why a simulation would need more than one narrow kind of Blackjack beginning.
Identify the situation dimensions: A Blackjack situation is organized by the player's sum, the dealer's showing card, and whether the player has a usable ace.
Identify the coverage risk: If starts were restricted, some combinations of these dimensions might receive little or no attention.
Apply exploring starts: The simulated episode can begin by selecting the dealer's cards, the player's sum, and the usable-ace condition at random with equal probability.
Connect coverage to improvement: Different situations can then contribute observed results to action values, giving the policy a basis for choosing actions across the described possibilities.
Exploring starts turn the beginning of each simulated game into a coverage mechanism rather than allowing the simulations to focus only on a narrow set of situations.
Values Behind the Policy
The policy is organized around the information available in a Blackjack situation: the player's sum, the dealer's showing card, and whether the player has a usable ace. For each such situation, the policy identifies an action such as hit or stick.
Action values describe choices made in particular situations. They associate a situation with an action, such as hitting or sticking, and represent the value used when comparing those choices. The state-value function is obtained from the action values. In other words, the value of a situation is derived from the action-value information available for the actions under the policy, rather than being an unrelated estimate.
Tracing One Improvement Cycle
A conceptual Monte Carlo ES cycle
Trace what happens when one simulated Blackjack episode begins from an exploring start.
Begin with the starting policy: The example begins with a policy that sticks on 20 or 21. This is only the starting policy, not the final reported result.
Arrange the episode start: The dealer's cards, the player's sum, and the usable-ace condition are selected at random with equal probability so that the situation is eligible for exploration.
Play the complete episode: The simulated game produces an observed result that can be used to evaluate the action choices encountered.
Update action values: The observed results contribute to action-value estimates for situation-action choices, including choices such as hit or stick.
Improve the policy: The updated action-value information is used to search for stronger action choices. Repeating this process produces the learned policy.
Monte Carlo ES converts complete simulated games into policy improvement by combining broad starts, observed results, action-value updates, and revised action choices.
| Policy comparison | What the source reports |
|---|---|
| Overall relationship | The Monte Carlo ES policy is similar to Thorp's basic strategy. |
| Reported difference | A leftmost notch in the policy for states with a usable ace is absent from Thorp's strategy. |
| Explanation of the difference | The source does not establish why this difference occurs. |
| Status of the learned policy | The policy is reported as optimal for the particular version of Blackjack described. |
Check Your Understanding
Explain why Monte Carlo ES uses exploring starts instead of relying only on the ordinary beginnings of simulated games. Then describe the sequence from an episode's observed result to an improved policy.
Hints
- Mention the player's sum, the dealer's cards, and the usable-ace condition.
- Distinguish action values from the state-value function.
- Remember that the starting stick-on-20-or-21 policy is not the final learned policy.
Treating Monte Carlo ES as only an evaluation method.
The method uses simulated games to search for a better policy, not merely to evaluate one fixed strategy.
Fix:
Track both stages: observed episode results update action values, and those values support policy improvement.Thinking that exploring starts mean randomizing only the first action.
The source describes random selection of the dealer's cards, the player's sum, and the usable-ace condition.
Fix:
Treat exploring starts as a coverage arrangement for Blackjack situations, with the listed elements selected at random with equal probability.Confusing an action value with a state value.
Action values describe state-action choices, while the state-value function is obtained from those action values.
Fix:
Keep the situation and the chosen action conceptually separate, then derive the state value from the relevant action-value information.Assuming that the starting stick-on-20-or-21 policy is the learned result.
The source describes it as the starting policy and reports that Monte Carlo ES later produces an optimal policy for the described game.
Fix:
Label the stick-on-20-or-21 rule as the initial policy and the Monte Carlo ES result as the learned policy.Claiming a reason for the difference from Thorp's strategy.
The source identifies the discrepancy but does not establish its reason.
Fix:
Report the difference without adding an unsupported cause.
Key Takeaways
- Monte Carlo ES uses complete simulated Blackjack games to search for a stronger policy, not merely to evaluate one fixed strategy.
- Exploring starts provide coverage by randomly selecting the dealer's cards, the player's sum, and the usable-ace condition with equal probability.
- The Blackjack example begins with a stick-on-20-or-21 policy; the supplied source does not state numerical initial action-value estimates.
- Action values describe state-action choices, and the state-value function is derived from those action values.
- The learned policy is reported as optimal for the described game and is similar to Thorp's basic strategy, with the source noting one difference involving a leftmost notch for usable-ace states.
Key Takeaways
- Monte Carlo ES turns simulated Blackjack episodes into policy improvement.
- Exploring starts ensure that different combinations of player sum, dealer cards, and usable-ace condition can be encountered.
- The example starts with a stick-on-20-or-21 policy; the supplied source does not specify numerical initial action-value estimates.
- State values are obtained from action values for the available choices.
- The learned policy is similar to Thorp's basic strategy, with one reported difference for usable-ace states.