Concepts / Representing Policies for Continuous Actions

Representing Policies for Continuous Actions

A normal density describes how probability is distributed across possible real-valued actions.

  • Programming

Choosing Actions from a Continuum

A policy for a continuous action cannot treat its action space like a short list of separate choices. A real-valued scalar action may take values across a continuous range. A normal density provides a way to represent how probability is distributed across those possible action values. The policy is represented by the complete density, not by a separate probability assigned to each exact value.

The central idea is to separate the policy representation from a probability query. The representation is the full normal density. A probability query asks about an interval of action values and uses the area under that density across the interval.

represented byassigns density acrosssupports samplingReal-valued actionspacepossible scalar actionsNormal densityprobability distributionDensity valuesacross action valuesAction sampleone real-valued action
How does a normal density assign likelihood across possible action values and how is a specific action sampled from it?

Density Is Not Point Probability

The value p(x) is a density at the action value x. It is not the same as the probability of taking exactly x. This distinction matters because a continuous-action policy is described by a density over real-valued actions rather than by separate probabilities for every exact action value.

Interpreting a Density Query

Suppose a policy is represented by a normal density over a real-valued scalar action. What is the difference between asking for p(x) at one action value and asking for the probability that the action lies within an interval?

Point query: The value p(x) tells us the density at the selected action value x. It describes the height of the density there, not the probability of taking exactly that value.

Interval query: To obtain probability for a range of action values, consider the area under the density across that range.

Representation: The normal density itself is the policy representation. The interval area is a probability query made using that representation.

A density value at one point and probability over an interval are different quantities. Probability comes from area across an action range.

evaluatesmeasuresAction value xone pointp(x)density at xAction intervalrange of valuesArea under densityprobability over interval
Why does the height of a normal density at one exact action value not equal the probability of taking that action, and how is probability obtained over an interval?

Mean and Spread Shape the Policy

A normal density is selected by its mean and standard deviation. Changing the mean produces a different density centered in a different location. Changing the standard deviation produces a different spread of density. These two values therefore act as policy parameters: together, they select which normal density represents the action policy.

Comparing Policy Parameters

Consider two versions of a normal-density policy. In the first version, the mean changes while the standard deviation stays fixed. In the second version, the standard deviation changes while the mean stays fixed. What aspect of the represented policy changes in each case?

Change the mean: The normal density is shifted to a different central location. The policy therefore represents a different distribution of action values.

Change the standard deviation: The spread of the normal density changes. The policy again represents a different distribution of action values, now because its standard deviation is different.

Keep both parameters: With a particular mean and standard deviation, the policy selects one particular normal density from the possible normal densities.

The mean controls the density's central location, while the standard deviation controls its spread. Both parameters help determine the policy's action distribution.

selects centercan increase spreadcan decrease spreadMeancentral locationShifted densitydifferent centerStandard deviationspreadWider densitygreater spreadNarrower densitysmaller spread
What changes in the action distribution when the mean shifts or the standard deviation increases or decreases?

A Normal Density as a Policy

For a continuous scalar action, the policy can be defined by a normal probability density over the possible action values. The density assigns different density values across that action space. A specific action can then be sampled from the represented distribution. This gives the policy a way to describe both where actions are centered and how broadly they are distributed.

represented bydescribessupports samplingScalar actionvaluesreal-valued possibilitiesNormal densitypolicy representationDistributed densityacross action valuesSampled actionone scalar value
How does a normal density assign likelihood across possible action values and how is a specific action sampled from it?

The policy is not a table containing one probability for every real-valued action. It is a normal density that describes how probability is distributed across the action space.

Function Approximators Produce the Parameters

In a continuous-action policy, the mean and standard deviation can come from parametric function approximators. The current state or observation is used by the approximator, and its outputs provide the parameters of the normal density. Those parameters then determine which normal density represents the policy for that situation.

Tracing One Policy Representation

Trace the information flow from a current observation to a continuous-action policy.

Start with the observation: The policy receives the current state or observation as the situation for which an action distribution is needed.

Apply the function approximator: A parametric function approximator processes that situation and provides the policy parameters.

Read the two parameters: The relevant outputs are the mean and standard deviation of the normal density.

Represent the action policy: Those parameters select the normal density that describes how probability is distributed across possible real-valued scalar actions.

Observation to function approximator to mean and standard deviation to normal density over scalar actions.

inputproducesproduceshelps determinehelps determineState orobservationcurrent situationFunction approximatorparametric mappingMeanpolicy parameterNormal densityaction policyStandard deviationpolicy parameter
How does the current state or observation flow through a function approximator to produce the policy's mean and standard deviation?
outputsoutputssets centersets spreadsupports samplingPolicy functionparametric approximatorMeancenterNormal densitypolicy representationReal-valued actionsampled valueStandard deviationspread
What do the two outputs of the policy function represent, and how do they determine the center and spread of sampled actions?

Mistakes in Reading Continuous Policies

  • Treating p(x) as the probability of taking exactly x.

    The source defines p(x) as a density. Probability is obtained from the area under the density across an interval.

    Fix: Use p(x) to describe density at a point, and use area across an action range when asking for probability.

  • Representing a continuous scalar policy as separate probabilities for every exact action value.

    The policy is represented by a normal probability density over the action values.

    Fix: Think of the full density as the policy representation.

  • Ignoring either the mean or the standard deviation.

    Changing either the mean or the standard deviation produces a different normal density.

    Fix: Treat both values as policy parameters: the mean determines central location and the standard deviation determines spread.

  • Confusing policy parameters with the action itself.

    The mean and standard deviation select the normal density; the action is a real-valued scalar sampled from that represented distribution.

    Fix: Keep the sequence clear: parameters define the density, and the density represents the action policy.

Check Your Understanding

MEDIUM

A function approximator receives the current observation and provides a mean and a standard deviation. Explain what these outputs determine, what the normal density represents, and how you would obtain probability for a range of possible scalar actions.

Hints
  • Identify the role of each output.
  • Distinguish the density at one point from area across an interval.
  • Describe the normal density as the policy representation.
  1. A continuous scalar action is represented with a normal density over real-valued action values. The value p(x) is density at one point, not the probability of taking exactly x. Probability over an action range comes from the area under the density across that range. The mean determines the density's central location, while the standard deviation determines its spread. Parametric function approximators can provide these two policy parameters from the current state or observation.

Key Takeaways

  • A normal density describes how probability is distributed across possible real-valued scalar actions.
  • The density value p(x) at one point is not the probability of taking exactly x.
  • Probability over an action interval is obtained from the area under the density across that interval.
  • The mean and standard deviation select the normal density by determining its central location and spread.
  • A parametric function approximator can use the current state or observation to provide the policy's mean and standard deviation.