Concepts / Gradient Ascent for Policy Optimization

Gradient Ascent for Policy Optimization

The policy gradient theorem links performance changes to the policy-weight vector θ.

  • Programming

Why the Theorem Matters

A policy is controlled by a vector of weights, written as θ. Changing these weights can change the policy, and changing the policy can change its performance. The policy gradient theorem describes how performance changes with respect to those policy weights. Its gradient is a column vector containing the partial derivatives with respect to the components of θ.

The central quantity is not a derivative with respect to an individual state or an arbitrary list of state probabilities. It is the change in policy performance described as a function of the policy-weight vector θ. The theorem connects that performance change to the direction in which the policy weights influence the policy.

controlsinfluencesdescribed with respect to θcomponentsθpolicy weightsπpolicy behaviorPolicy gradientpartial derivatives withrespect to θPerformancepolicy objective
How does the policy gradient theorem connect performance changes to the policy-weight vector?

Following Visits Through Episodes

To understand dπ(s), imagine starting every episode at the same starting state s0. The policy π chooses actions, while the MDP dynamics determine how each episode develops. Because both the policy and the dynamics are involved, the resulting episodes can vary. For one particular state s, count how many time steps each episode spends in that state. The quantity dπ(s) is the expected value of that count across the randomly generated episodes.

followed fromdeveloped undercontributes behaviorcontributes developmentcountaverage expected counts0episode startπaction choicesEpisodesrandomly generatedVisits to stime-step countsdπ(s)expected countMDP dynamicsepisode development
How do action choices from the policy and transition dynamics combine to determine how often a state is visited?

Counting Visits to One State

Consider three generated episodes that all begin at s0. In these illustrative episodes, count the time steps spent in state s.

Episode 1: The sequence visits s twice, so its count for state s is 2.

Episode 2: The sequence visits s once, so its count for state s is 1.

Episode 3: The sequence does not visit s, so its count for state s is 0.

Expected count: Across these three illustrative episodes, the average count is (2 + 1 + 0) divided by 3, which is 1.

In this illustration, dπ(s) is 1. The value represents an expected number of time-step visits, not the visit count of one particular episode.

Reading dπ(s)

dπ(s) is the expected number of time-step visits to state s across episodes that start at s0 and follow both the policy π and the MDP dynamics.

The subscript π indicates that the visit count depends on the policy behavior. However, the policy is not the only influence. The MDP dynamics also determine how an episode unfolds after actions are chosen. Therefore, dπ(s) summarizes the combined effect of policy behavior and MDP dynamics on how often state s is encountered on average.

countcountcountcombinecombinecombineEpisode 1s0 → s → s2visits to sdπ(s)expected countEpisode 2s0 → s1visit to sEpisode 3s00visits to s
What does dπ(s) contain, and how do separate episode sequences contribute to the expected count for state s?

Separating the Gradient Targets

The policy gradient theorem differentiates performance with respect to the policy-weight vector θ. The resulting gradient is a column vector whose entries are partial derivatives with respect to the individual components of θ.

The theorem also uses dπ(s), because the expected state visits help describe the episodes produced by the policy and the MDP. But the theorem's important contribution is that the expression does not require taking the derivative of the state distribution. In other words, dπ(s) appears as part of the theorem's performance-gradient expression; it is not introduced as a separate gradient target in the theorem's policy-weight derivative.

differentiate with respect tocomponents determineappears in expressionseparate derivativePerformancequantity that changesθpolicy weights∇θ performancecolumn vectordπ(s)expected state visitsDerivative of dπ(s)not required by the theorem
Which parts of the performance expression does the theorem differentiate with respect to θ, and which state-distribution terms are not treated as separate gradient targets?
QuantityRole in the theorem
PerformanceThe quantity whose change is described
θThe policy-weight vector with respect to which the gradient is taken
Components of θThe variables associated with the partial derivatives in the gradient column vector
dπ(s)The expected time-step count for visits to state s
Derivative of dπ(s)Not required by the theorem's expression

What θ Controls

Changing One Policy Weight

Suppose a policy is controlled by a vector θ with several components. Consider changing one component while leaving the other components unchanged.

Weight change: The selected component of θ changes. Because θ controls the policy, this change can change the policy's behavior.

Episode consequences: A changed policy can produce different action choices. Together with the MDP dynamics, those choices can change how episodes develop and how often states are visited.

Performance consequence: Because the policy may have changed, its performance may also change.

Gradient interpretation: The corresponding component of the policy gradient describes performance change with respect to that component of θ. All such partial derivatives form the gradient column vector.

θ is the coordinate system for the policy-weight gradient: its components are the variables whose influence on performance is described.

This does not mean that each component of θ directly equals a state visit count. The weights control the policy, the policy helps determine behavior, and the policy together with the MDP dynamics determines the episodes summarized by dπ(s). The gradient then describes performance changes with respect to the original policy-weight components.

Mistakes to Avoid

  • Treating dπ(s) as the count from one episode

    dπ(s) is the expected count across randomly generated episodes that start at s0 and follow the policy and MDP dynamics.

    Fix: Count visits in each episode, then interpret dπ(s) as the expected value of those counts.

  • Attributing dπ(s) only to the policy

    The MDP dynamics also determine how the episode develops after actions are chosen.

    Fix: Remember that policy behavior and MDP dynamics jointly determine the episodes used to define dπ(s).

  • Thinking the theorem takes a separate gradient with respect to dπ(s)

    The theorem's important contribution is that its expression does not require taking the derivative of the state distribution.

    Fix: Identify θ as the variable of differentiation and dπ(s) as the expected state-visit quantity used in the expression.

  • Describing the gradient as one number

    The gradient is a column vector containing partial derivatives with respect to the components of θ.

    Fix: Connect each entry of the gradient to a component of the policy-weight vector.

When reading a policy-gradient expression, label each quantity by its role: performance is the quantity whose change is described, θ supplies the policy-weight variables, and dπ(s) summarizes expected state visits under the policy and MDP dynamics.

Practice Check

MEDIUM

An episodic process starts every episode at s0. The policy π and the MDP dynamics generate several episodes. In one episode, state s is visited three times; in another, it is visited zero times. Explain what additional information is needed to determine the expected count dπ(s), and identify the quantity with respect to which the policy gradient theorem differentiates performance.

Hints
  • dπ(s) is based on expected counts across the randomly generated episodes, not just one or two selected counts.
  • The gradient is taken with respect to the components of the policy-weight vector θ.

Key Takeaways

  1. The policy gradient theorem describes how policy performance changes with respect to the policy-weight vector θ.
  2. The gradient is a column vector of partial derivatives with respect to the components of θ.
  3. dπ(s) is the expected number of time-step visits to state s across episodes beginning at s0.
  4. The policy and the MDP dynamics jointly determine the episodes from which dπ(s) is defined.
  5. The theorem uses the state distribution without requiring a separate derivative of that state distribution.

Key Takeaways

  • The policy gradient theorem connects changes in performance to the policy-weight vector θ.
  • Its gradient contains partial derivatives for the components of θ.
  • dπ(s) records the expected time-step count for visits to state s across episodes starting at s0.
  • Both policy behavior and MDP dynamics determine the episodes summarized by dπ(s).
  • The theorem does not require taking the derivative of the state distribution as a separate gradient target.