Concepts / Policy Parameterization for Continuous Actions

Policy Parameterization for Continuous Actions

A normal density describes how probability is distributed across possible real-valued actions.

  • Programming

From State to Action Distribution

A policy for a continuous scalar action does not assign a separate probability to every possible real number. Instead, it defines a normal probability density over the possible action values. The density is parameterized by a mean and a standard deviation, and those parameters can be produced by parametric function approximators.

inputprovidesprovidesparameterizesparameterizesStateobserved situationFunction approximatorpolicy parameterizerMeannormal-density parameterNormal densitypolicy over scalar actionsStandard deviationnormal-density parameter
How does an observed state flow through a function approximator to produce the parameters of a normal policy?

Tracing a Continuous Policy

Suppose the action is a real-valued scalar, such as a single control setting. The policy represents possible values of that action with a normal density. The full density p(x) is the policy representation: it describes how probability is distributed across possible action values.

A Scalar Action Policy

A policy must represent possible real-valued actions for one observed state. How does the normal-density representation answer an action question?

Parameterize: A parametric function approximator receives the state and provides a mean and a standard deviation.

Represent: Those two parameters select a normal density over the possible scalar action values.

Query: To ask about a range of actions, consider the area under the selected density across that range.

The policy is the selected normal density, while a probability question concerns an interval and the area under the density over that interval.

selectdistributes probability acrossMean and standarddeviationpolicy parametersNormal densityp(x)Scalar actionsreal-valued possibilities
How does a normal density turn policy parameters into probabilities for different possible scalar actions?

Density Versus Probability

The value p(x) at one action value is a density, not the probability of taking exactly that action value. A density describes how probability is distributed locally across the action axis. Probability is obtained over a range of action values: it is the area under the density across that interval.

givesgivesxdensity p(x)Densityheight at one valueAction intervalarea under p(x)Probabilityarea across a range
What is the difference between the height of a density at one action value and the probability contained across an action interval?

Reading a Policy Query Correctly

A learner looks at p(x) for one action value and calls that number the probability of taking exactly x. What needs to be corrected?

Identify the quantity: The value p(x) is a density at a point.

Identify the query: A probability query must refer to a range of action values rather than only one exact value.

Use the density: The probability for that range is represented by the area under the density across the range.

The point value is a density; the interval area is the probability.

Changing the Mean

The mean is one of the parameters that selects the normal density used by the policy. Changing the mean produces a different density and shifts where the density is centered along the action axis. In policy terms, changing the mean changes the central action region represented by the distribution.

mean changesMean Adensity centered at onelocationMean Bdensity centered at anotherlocation
How does changing the mean shift the normal density along the action axis?

Changing the Standard Deviation

The standard deviation is the second policy parameter that selects the normal density. Changing it produces a different density with a different spread. A larger spread distributes the density across a wider region of action values, while a smaller spread concentrates it more narrowly. Thus, the standard deviation controls how broadly the policy represents possible scalar actions.

standard deviation changesSmall standarddeviationnarrower densityLarge standarddeviationwider density
How does changing the standard deviation alter the spread and peak of a normal policy density?
ParameterWhat changing it doesPolicy interpretation
MeanChanges the location of the normal densityChanges the central action region
Standard deviationChanges the spread of the normal densityChanges how broadly possible actions are represented

The two parameters select different aspects of the normal policy density.

Function Approximators as Policy Parameterizers

For a continuous-action policy, parametric function approximators provide the mean and standard deviation of the normal density. The observed state is passed through the approximator, and its outputs determine which normal density represents the policy for that state. A different state can therefore lead to different parameter values and a different normal density.

This separates two jobs. The function approximator supplies the policy parameters. The normal density uses those parameters to represent how probability is distributed over real-valued scalar actions. The policy is not merely the mean by itself; it is the density selected by both the mean and the standard deviation.

Mistakes to Avoid

  • Treating p(x) as the probability of one exact action value.

    p(x) is a density at a point. Probability is associated with the area under the density across an interval.

    Fix: For a probability question, identify an action range and reason about the area under the density over that range.

  • Describing the policy only by its mean.

    The normal density is selected by both its mean and its standard deviation.

    Fix: Track both outputs from the parametric function approximator.

  • Confusing a parameter change with a probability query.

    Changing the mean produces a different density; an interval probability is then obtained from the area under that selected density.

    Fix: First identify the density parameters, then distinguish the resulting density from any interval-based probability question.

Practice the Trace

MEDIUM

A function approximator receives a state and produces a mean and a standard deviation for a normal policy. Explain the full chain from the state to a probability question about a range of scalar actions. Then state why the density value at one exact action is not itself that action's probability.

Hints
  • Name the two outputs of the function approximator.
  • Explain what those outputs select.
  • Distinguish a point value of the density from an interval area.

What do you think happens?

If the mean changes while the standard deviation stays fixed, what aspect of the normal policy density changes?

  • Its location along the action axis
  • Only the name of the policy
  • Whether p(x) is a density
  • The fact that actions are real-valued
Reveal answer

Answer: Its location along the action axis

The mean is a policy parameter that changes which normal density is selected and shifts the density's central location.

Key Takeaways

  1. A continuous scalar-action policy can be represented by a normal probability density over real-valued actions.
  2. The value p(x) at one point is a density, not the probability of one exact action value.
  3. Probability over an action range comes from the area under the density across that interval.
  4. The mean changes the location of the normal density, while the standard deviation changes its spread.
  5. Parametric function approximators can receive a state and provide the mean and standard deviation that select the policy density.

Key Takeaways

  • A normal density represents how probability is distributed across possible real-valued scalar actions.
  • Point density and interval probability are different quantities.
  • The mean and standard deviation parameterize the normal policy density.
  • Parametric function approximators provide those policy parameters from the observed state.