Deep Neural Network Function Approximation
A parameterized policy uses weights θ to determine action probabilities.
From Choices to Probabilities
A policy must ultimately choose actions, but a policy-gradient method needs more than a final action choice. It needs a parameterized policy whose action probabilities depend on weights θ. The policy can first assign numerical preferences to available state-action pairs and then convert those preferences into probabilities.
The central separation is between producing preferences and converting preferences into action probabilities.
Organizing State-Action Preferences
For a particular state s, consider each available action a. A preference function h(s, a, θ) assigns a numerical preference to that state-action pair. These numbers are not required to be action values of the kind used by an action-value method. They are inputs to the policy's probability rule: actions with higher preferences receive higher selection probabilities.
Assigning Preferences Before Probabilities
For one state s, suppose three available actions receive the preferences h(s, a₁, θ) = 4, h(s, a₂, θ) = 1, and h(s, a₃, θ) = -2.
List the pairs: The policy considers the same state s paired with each available action.
Attach one number to each pair: The preference function assigns 4 to a₁, 1 to a₂, and -2 to a₃.
Compare the preferences: Action a₁ has the highest preference, followed by a₂, while a₃ has the lowest preference.
Apply the probability rule: The preference values are then passed to the exponential softmax rule, which gives higher probability to the higher-preference actions and normalizes the results.
The preference stage produces an ordered set of numerical inputs for the policy's probability stage; it does not by itself state the final probabilities.
Normalizing with Exponential Softmax
π(a | s, θ) = exp(h(s, a, θ)) / Σb∈A(s) exp(h(s, b, θ))
The exponential softmax rule preserves the preference ordering: a higher preference produces a higher probability. The denominator includes all available actions, so the probabilities are normalized across the action set for the state.
A Two-Action Softmax Calculation
For one state, suppose two actions have preferences h(s, a₁, θ) = ln(2) and h(s, a₂, θ) = 0.
Exponentiate each preference: The exponential preferences are exp(ln(2)) = 2 and exp(0) = 1.
Build the normalization sum: The denominator is 2 + 1 = 3.
Divide each exponential preference: Action a₁ receives 2/3 and action a₂ receives 1/3.
π(a₁ | s, θ) = 2/3 and π(a₂ | s, θ) = 1/3. The higher-preference action receives the higher probability, and the probabilities sum to 1.
Choosing a Preference Function
The softmax rule does not dictate how the preference values must be produced. The preference function h(s, a, θ) can be parameterized arbitrarily. A deep neural network is one possibility, with θ representing its vector of connection weights. A function that is linear in feature vectors φ(s, a) is another possibility. These choices change the computation of preferences, while the same softmax rule can convert those preferences into action probabilities.
The important design boundary is therefore clear: the parameterization determines how θ produces h(s, a, θ), and the softmax determines how h becomes a policy. Replacing a linear feature function with a deep neural network changes the preference-producing mechanism, not the basic preference-to-probability rule.
Why Differentiability Matters
Policy gradient methods adjust the policy through its weights. For that adjustment to be expressed as a gradient, the policy must be differentiable with respect to those weights. The required gradient must exist and must always be finite.
Common Reasoning Errors
Treating a preference as an action probability
Preferences are numerical inputs to the probability rule. They are not yet normalized action probabilities.
Fix:
Apply exponential softmax across all available actions before interpreting the results as probabilities.Normalizing one action independently
The softmax denominator uses the exponential preferences of all available actions for the state.
Fix:
Compute the shared normalization sum over the complete available-action set.Assuming the preference function must be a neural network
The preference function can be parameterized arbitrarily; a deep neural network and a linear feature function are both identified possibilities.
Fix:
Distinguish the mechanism that produces preferences from the softmax rule that converts them into probabilities.Ignoring differentiability
Policy gradient methods require a policy gradient that exists and is always finite.
Fix:
Check the differentiability requirement with respect to the policy weights.Confusing preferences with action values
A policy does not have to assign the same kind of value to actions that an action-value method does.
Fix:
Treat preferences as quantities used to rank actions and determine their selection probabilities.
Check Your Understanding
A policy uses a preference function h(s, a, θ) for three actions in one state. Explain the complete path from policy weights θ to the final probability of selecting one action. Your explanation should identify the preference stage, the exponential softmax stage, and the differentiability requirement.
Hints
- Start with how θ determines numerical preferences for state-action pairs.
- State what the exponential softmax numerator and denominator contain.
- Explain why the policy must have a gradient with respect to θ.
What do you think happens?
Suppose two actions have preferences ln(2) and 0. Before calculating, which action receives the larger probability, and what is the probability of each action after softmax?
Reveal answer
Answer: The first action receives 2/3 and the second receives 1/3.
Exponentiating the preferences gives 2 and 1. Their normalization sum is 3, so the probabilities are 2/3 and 1/3.
Key Takeaways
- A parameterized policy uses weights θ to determine action probabilities.
- The preference function h(s, a, θ) assigns numerical preferences to state-action pairs before probabilities are computed.
- Exponential softmax gives higher probability to higher-preference actions and normalizes the probabilities across the available actions.
- The preference function can be produced by a deep neural network, a linear function of feature vectors, or another parameterization.
- Policy gradient methods require the policy to be differentiable with respect to its weights, with a gradient that exists and is always finite.