Policy Parameterization
Softmax action selection over action preferences can approach deterministic behavior when the optimal action's preference is driven far above those of suboptimal actions.
From Preferences to Actions
A learning method can represent either the value of each possible action or the rule that selects actions. Policy parameterization focuses on the second object: it represents the policy directly. One way to do this is to assign each action a preference and use softmax action selection to convert those preferences into action probabilities.
Following the Preference Gap
Early in learning, several actions can have noticeable preferences. Softmax action selection can therefore produce a stochastic policy: more than one action may retain a meaningful chance of being selected. As learning drives the preference for the optimal action farther above the preferences of suboptimal actions, the probability of selecting the optimal action can approach one.
What do you think happens?
Suppose one action's preference keeps increasing while the preferences of the other actions remain far lower. What happens to the resulting policy?
Reveal answer
Answer: The preferred action is selected almost every time.
When the optimal action's preference is driven far above the preferences of suboptimal actions, softmax action selection can move its probability toward one. The behavior approaches deterministic action selection.
Two Objects to Approximate
An action-value method estimates how valuable each action is. A policy-approximation method represents the action-selection rule more directly. These methods can face the same decision problem while learning different objects, and the object being learned can have a simpler or more complicated structure depending on the problem.
If the policy is a simpler function than the action-value function for a particular problem, policy-based learning may learn faster and produce a better final policy. The reverse is also possible: if the action-value function has the simpler structure, an action-value method may be more convenient. There is no universal rule that one representation is always simpler.
Choosing the Simpler Representation
Two methods must solve the same decision problem. One represents action values, and the other represents the action-selection policy directly. How should the representation choice be evaluated?
Identify the learned object: The action-value method learns how valuable actions are. The policy method learns how actions should be selected.
Compare structural complexity: Ask which of those two objects has the more manageable structure in this particular problem.
Choose conditionally: If the policy is simpler, policy-based learning may be faster and more effective. If the action-value function is simpler, an action-value method may be the more convenient choice.
The advantage comes from matching the learning method to the simpler object, not from assuming that policies are always simpler.
When Randomness Is Optimal
Deterministic behavior is not always the best behavior. In an imperfect-information problem, consistently choosing one action can reveal a predictable pattern. The best policy may instead assign particular probabilities to different actions. Card games with imperfect information provide an example: optimal play can require choosing actions such as bluffing with specific probabilities.
This is a distinctive expressive advantage of policy approximation. Its target is the policy itself, so it can represent a stochastic optimal policy directly. Action-value methods are centered on selecting actions from value estimates and do not have a natural way to find a policy whose required behavior is a deliberate probability distribution.
A Policy for Imperfect Information
Consider an imperfect-information card-game situation in which predictable behavior can reveal a pattern. Compare always taking one action with a policy that assigns specific probabilities to two actions.
Test the deterministic choice: Always selecting one action creates a predictable pattern. In an imperfect-information problem, that predictability can be harmful.
Represent deliberate randomization: A policy can assign different actions particular probabilities, such as choosing when to bluff in Poker.
Match the representation to the goal: Because the policy representation directly describes action selection, it can express the required stochastic behavior itself.
When optimal behavior deliberately randomizes, directly approximating the policy can express the target behavior more naturally than estimating action values.
Where Policy Approximation Helps
- Policy approximation can move toward deterministic behavior by increasing the preference gap, rather than retaining a forced random-action probability.
- It may be advantageous when the policy has a simpler structure than the action-value function.
- It can represent an optimal policy that deliberately randomizes, which is important in some imperfect-information problems.
Common Reasoning Errors
Assuming that a stochastic policy is merely unfinished learning.
Early stochastic behavior can reflect learning progress, but stochastic behavior can also be the desired optimal behavior in an imperfect-information problem.
Fix:
Ask whether the problem rewards deliberate randomization before treating stochastic behavior as a defect.Treating softmax action selection as if it retained epsilon-greedy exploration by definition.
Epsilon-greedy retains an epsilon probability of a random action, whereas softmax over action preferences can approach deterministic behavior as the preference gap grows.
Fix:
Separate forced random exploration from probabilities produced by relative action preferences.Assuming the policy is always simpler than the action-value function.
Some problems have simpler action-value functions instead.
Fix:
Compare the complexity of the policy and the action-value function for the particular problem.Assuming action-value methods and policy methods learn the same object.
Action-value methods estimate action values, while policy approximation represents the action-selection rule more directly.
Fix:
Name the learned object before evaluating the method.
Check Your Understanding
A problem has an optimal policy that deliberately chooses two actions with particular probabilities. Explain why directly approximating the policy may be a better fit than relying on action selection centered on estimated action values. Then describe a different problem property that could make action-value approximation preferable.
Hints
- Focus on which object each method represents.
- Consider whether the desired behavior is deterministic or stochastic.
- Remember that the simpler object can differ from one problem to another.
- Softmax action selection converts action preferences into a policy. When one preference is driven far above the others, the policy can approach deterministic selection. Policy approximation can be useful when the policy is simpler than the action-value function or when the optimal behavior is deliberately stochastic. The correct choice depends on the structure of the problem.
Key Takeaways
- Policy parameterization represents the action-selection rule directly, often by assigning preferences to actions and converting them into probabilities with softmax selection.
- Increasing the preference gap can make softmax behavior approach deterministic action selection.
- A policy and an action-value function are different learned objects, and either one may be simpler depending on the problem.
- Policy approximation can directly represent optimal stochastic behavior, such as deliberate randomization in imperfect-information problems.
- Policy approximation is advantageous when its representation fits the problem better; action-value methods may be preferable when the action-value function is simpler.