Concepts / Stochastic Policies

Stochastic Policies

Softmax action selection over action preferences can approach deterministic behavior when the optimal action's preference is driven far above those of suboptimal actions.

  • Programming

From Preference to Behavior

A policy is an action-selection rule. It may choose one action consistently, or it may assign different probabilities to several actions. Softmax action selection over action preferences provides a way for the same policy representation to behave stochastically early in learning and to approach deterministic behavior when one action's preference becomes much higher than the others.

What do you think happens?

Suppose one action's preference rises far above the preferences of all other actions. What should happen to the policy's action probabilities?

  • The action probabilities remain equally divided
  • The dominant action becomes increasingly likely
  • Every action becomes equally unlikely
Reveal answer

Answer: The dominant action becomes increasingly likely.

With softmax action selection, driving the optimal action's preference far above the preferences of suboptimal actions makes the policy approach choosing that action almost every time.

Softmax Preference Trace

softmax selectionsoftmax selectionAction preferencesseveral actions remaincompetitiveOptimal actionpreferencefar above suboptimalpreferencesStochastic policyseveral actions retainnoticeable probabilitiesOptimal actionchosen almost every time
How do action preferences translate into selection probabilities, and what happens as one preference rises far above the others?

Softmax action selection converts relative action preferences into selection probabilities. When several actions have noticeable preferences, the policy can retain noticeable probabilities for several actions and therefore behave stochastically. If the preference for the optimal action is driven infinitely higher than the preferences for suboptimal actions, when the parameterization permits this, its probability can move toward one. The policy then approaches deterministic behavior: it chooses the optimal action almost every time.

A Policy Versus Action Values

representsrepresentsPolicyaction-selection ruleAction-valuefunctionvalue for each state-actionpairPolicy structuremay be simplerValue structuremay be more complex
How can the same decision problem lead one method to represent an action-selection rule and another to represent values for every state-action pair?

Two methods can face the same decision problem while representing different objects. An action-value method estimates how valuable each action is for each state-action pair. A policy-approximation method represents the action-selection rule more directly. These objects do not have to have the same complexity. In some problems, a simple policy may correspond to an action-value function with a more complicated pattern of values.

Choosing the Easier Object to Represent

Imagine a decision problem in which the useful behavior can be described by a relatively simple action-selection rule, while the values of individual actions vary in a more complicated way across state-action pairs.

Identify the targets: An action-value method targets the values of individual actions. A policy method targets the action-selection rule itself.

Compare the structures: The policy may have the more manageable structure, even though the action-value function must describe many complicated values.

Choose the approximation target: A policy-based method may learn faster and produce a better final policy when the policy is simpler than the action-value function.

The simpler object is not always the action-value function. When the policy is simpler, directly approximating the policy can be advantageous.

When Randomization Is Optimal

requiresselectsselectsImperfectinformationdecision problemAction onespecific probabilityStochastic policyaction probabilitiesAction twospecific probability
How does a policy represent a decision problem in which optimal behavior requires different actions with specific probabilities?

In imperfect-information problems, deterministic behavior can reveal a predictable pattern. The best policy may instead assign different actions specific probabilities. The source gives card games with imperfect information as an example: optimal play can require doing two different things with particular probabilities, such as choosing when to bluff in Poker.

A Deliberately Stochastic Card-Game Policy

Consider a card-game situation in which an opponent could exploit a completely predictable action choice. The desired behavior is to use two actions with particular probabilities rather than always using one action.

Recognize the information problem: Because the situation contains imperfect information, consistently choosing one action can reveal a predictable pattern.

Represent the desired behavior: A stochastic policy can directly assign probabilities to the available actions.

Preserve randomization: The randomization is part of the target behavior, not merely temporary exploration that should disappear as learning continues.

Policy approximation is suitable because its target is the policy itself, including a policy that deliberately randomizes.

Why Policy Approximation Can Help

represents directlysupportsPolicyapproximationrepresents action-selectionruleAction-value methodestimates action valuesAction probabilitiesdirect policy outputAction selectionbased on value estimates
What is different in the information and computation flow when a method represents action probabilities directly instead of selecting actions from estimated values?

Policy approximation has three possible advantages over action-value methods. First, softmax preferences can move toward deterministic behavior without retaining forced random exploration. The source contrasts this with epsilon-greedy selection, which retains an epsilon probability of a random action. Second, the policy may be simpler than the action-value function, allowing policy-based learning to learn faster and produce a better final policy. Third, policy approximation can represent an optimal policy that deliberately randomizes, because its target is the policy itself.

QuestionPolicy approximationAction-value method
What is represented?The action-selection rule directlyThe value of each action for each state-action pair
Can the representation approach deterministic behavior?Yes, through preferences whose relative differences become very largeIts usual action-selection approach is centered on selecting actions from value estimates
Can the target remain deliberately stochastic?Yes, it can represent an optimal policy that assigns specific probabilitiesIt does not have a natural way to find such policies through its usual value-centered action selection
When may it be simpler?When the policy has a more manageable structureWhen the action-value function has a more manageable structure

Mistakes About Stochastic Policies

  • Assuming that stochastic behavior is always temporary exploration.

    In imperfect-information problems, the best policy may deliberately assign different actions specific probabilities.

    Fix: Ask whether randomization is part of the optimal policy or merely a learning behavior that should eventually fade.

  • Assuming that softmax automatically has the same exploration behavior as epsilon-greedy selection.

    Epsilon-greedy explicitly retains an epsilon probability of a random action, while softmax over action values does not automatically provide that same route to determinism.

    Fix: Distinguish the explicit random-action probability in epsilon-greedy selection from the probability distribution produced by softmax preferences.

  • Assuming that action-value functions are always simpler than policies.

    Some problems have a simpler policy and a more complicated action-value function.

    Fix: Compare the structure of the policy with the structure of the action-value function before choosing the approximation target.

  • Assuming that policy approximation is always superior.

    The source states that the reverse can occur: simpler action-value functions can make action-value methods more convenient.

    Fix: Treat the advantage as problem-dependent rather than universal.

Apply the Distinction

MEDIUM

A decision problem has two properties: its optimal behavior requires two actions to occur with particular probabilities, and the policy describing that behavior appears simpler than the action-value function. Which learning target is likely to be attractive, and why?

Hints
  • Identify whether the optimal behavior is deterministic or stochastic.
  • Identify which object has the more manageable structure.
  • Remember that policy approximation targets the action-selection rule directly.

Practice Answer

Use the two properties in the practice situation to select a learning target.

Check the target behavior: The optimal behavior is stochastic, so a method that can represent action probabilities directly is relevant.

Check the representation complexity: The policy is described as simpler than the action-value function, which favors approximating the policy.

Combine the reasons: Both the desired stochastic behavior and the simpler policy structure support policy approximation.

Policy approximation is likely to be attractive because it can represent the stochastic optimal policy directly and may learn more effectively when the policy is simpler.

Key Takeaways

  1. Softmax action selection can behave stochastically when several action preferences remain competitive.
  2. When the optimal action's preference rises far above the others, its selection probability can approach one, producing nearly deterministic behavior.
  3. A policy and an action-value function are different approximation targets, and either one may have the simpler structure depending on the problem.
  4. Policy approximation is especially suitable when the optimal behavior deliberately randomizes, as can occur in imperfect-information card games.
  5. Policy approximation may avoid forced random exploration, exploit a simpler policy structure, and represent stochastic optimal policies directly.

Key Takeaways

  • Softmax converts relative action preferences into action-selection probabilities.
  • A dominant optimal-action preference can make the resulting policy approach deterministic behavior.
  • The policy may be simpler or more complex than the action-value function; the better approximation target depends on the problem.
  • Policy approximation directly represents stochastic behavior when randomization is part of optimal play.
  • Its potential advantages include avoiding forced random exploration, learning a simpler target, and expressing deliberately randomized policies.