Policy Parameterization for Continuous Actions
A normal density describes how probability is distributed across possible real-valued actions.
From State to Action Distribution
A policy for a continuous scalar action does not assign a separate probability to every possible real number. Instead, it defines a normal probability density over the possible action values. The density is parameterized by a mean and a standard deviation, and those parameters can be produced by parametric function approximators.
Tracing a Continuous Policy
Suppose the action is a real-valued scalar, such as a single control setting. The policy represents possible values of that action with a normal density. The full density p(x) is the policy representation: it describes how probability is distributed across possible action values.
A Scalar Action Policy
A policy must represent possible real-valued actions for one observed state. How does the normal-density representation answer an action question?
Parameterize: A parametric function approximator receives the state and provides a mean and a standard deviation.
Represent: Those two parameters select a normal density over the possible scalar action values.
Query: To ask about a range of actions, consider the area under the selected density across that range.
The policy is the selected normal density, while a probability question concerns an interval and the area under the density over that interval.
Density Versus Probability
The value p(x) at one action value is a density, not the probability of taking exactly that action value. A density describes how probability is distributed locally across the action axis. Probability is obtained over a range of action values: it is the area under the density across that interval.
Reading a Policy Query Correctly
A learner looks at p(x) for one action value and calls that number the probability of taking exactly x. What needs to be corrected?
Identify the quantity: The value p(x) is a density at a point.
Identify the query: A probability query must refer to a range of action values rather than only one exact value.
Use the density: The probability for that range is represented by the area under the density across the range.
The point value is a density; the interval area is the probability.
Changing the Mean
The mean is one of the parameters that selects the normal density used by the policy. Changing the mean produces a different density and shifts where the density is centered along the action axis. In policy terms, changing the mean changes the central action region represented by the distribution.
Changing the Standard Deviation
The standard deviation is the second policy parameter that selects the normal density. Changing it produces a different density with a different spread. A larger spread distributes the density across a wider region of action values, while a smaller spread concentrates it more narrowly. Thus, the standard deviation controls how broadly the policy represents possible scalar actions.
| Parameter | What changing it does | Policy interpretation |
|---|---|---|
| Mean | Changes the location of the normal density | Changes the central action region |
| Standard deviation | Changes the spread of the normal density | Changes how broadly possible actions are represented |
The two parameters select different aspects of the normal policy density.
Function Approximators as Policy Parameterizers
For a continuous-action policy, parametric function approximators provide the mean and standard deviation of the normal density. The observed state is passed through the approximator, and its outputs determine which normal density represents the policy for that state. A different state can therefore lead to different parameter values and a different normal density.
This separates two jobs. The function approximator supplies the policy parameters. The normal density uses those parameters to represent how probability is distributed over real-valued scalar actions. The policy is not merely the mean by itself; it is the density selected by both the mean and the standard deviation.
Mistakes to Avoid
Treating p(x) as the probability of one exact action value.
p(x) is a density at a point. Probability is associated with the area under the density across an interval.
Fix:
For a probability question, identify an action range and reason about the area under the density over that range.Describing the policy only by its mean.
The normal density is selected by both its mean and its standard deviation.
Fix:
Track both outputs from the parametric function approximator.Confusing a parameter change with a probability query.
Changing the mean produces a different density; an interval probability is then obtained from the area under that selected density.
Fix:
First identify the density parameters, then distinguish the resulting density from any interval-based probability question.
Practice the Trace
A function approximator receives a state and produces a mean and a standard deviation for a normal policy. Explain the full chain from the state to a probability question about a range of scalar actions. Then state why the density value at one exact action is not itself that action's probability.
Hints
- Name the two outputs of the function approximator.
- Explain what those outputs select.
- Distinguish a point value of the density from an interval area.
What do you think happens?
If the mean changes while the standard deviation stays fixed, what aspect of the normal policy density changes?
Reveal answer
Answer: Its location along the action axis
The mean is a policy parameter that changes which normal density is selected and shifts the density's central location.
Key Takeaways
- A continuous scalar-action policy can be represented by a normal probability density over real-valued actions.
- The value p(x) at one point is a density, not the probability of one exact action value.
- Probability over an action range comes from the area under the density across that interval.
- The mean changes the location of the normal density, while the standard deviation changes its spread.
- Parametric function approximators can receive a state and provide the mean and standard deviation that select the policy density.
Key Takeaways
- A normal density represents how probability is distributed across possible real-valued scalar actions.
- Point density and interval probability are different quantities.
- The mean and standard deviation parameterize the normal policy density.
- Parametric function approximators provide those policy parameters from the observed state.