Concepts / Deep Neural Network Function Approximation

Deep Neural Network Function Approximation

A parameterized policy uses weights θ to determine action probabilities.

  • Programming

From Choices to Probabilities

A policy must ultimately choose actions, but a policy-gradient method needs more than a final action choice. It needs a parameterized policy whose action probabilities depend on weights θ. The policy can first assign numerical preferences to available state-action pairs and then convert those preferences into probabilities.

The central separation is between producing preferences and converting preferences into action probabilities.

determinesoftmax convertsassign selection likelihoodPolicy weights θparameterized policyPreferences h(s, a,θ)one value per state-actionpairAction probabilitiessoftmax distributionSelected actionpolicy choice
How does a change in policy weights flow through the policy computation to change the probability of selecting an action?

Organizing State-Action Preferences

For a particular state s, consider each available action a. A preference function h(s, a, θ) assigns a numerical preference to that state-action pair. These numbers are not required to be action values of the kind used by an action-value method. They are inputs to the policy's probability rule: actions with higher preferences receive higher selection probabilities.

paired withpaired withpaired withState ssame state for each pairAction a₁h(s, a₁, θ)Action a₂h(s, a₂, θ)Action a₃h(s, a₃, θ)
How are preference values organized for the available actions given a particular state?

Assigning Preferences Before Probabilities

For one state s, suppose three available actions receive the preferences h(s, a₁, θ) = 4, h(s, a₂, θ) = 1, and h(s, a₃, θ) = -2.

List the pairs: The policy considers the same state s paired with each available action.

Attach one number to each pair: The preference function assigns 4 to a₁, 1 to a₂, and -2 to a₃.

Compare the preferences: Action a₁ has the highest preference, followed by a₂, while a₃ has the lowest preference.

Apply the probability rule: The preference values are then passed to the exponential softmax rule, which gives higher probability to the higher-preference actions and normalizes the results.

The preference stage produces an ordered set of numerical inputs for the policy's probability stage; it does not by itself state the final probabilities.

Normalizing with Exponential Softmax

π(a | s, θ) = exp(h(s, a, θ)) / Σb∈A(s) exp(h(s, b, θ))

The exponential softmax rule preserves the preference ordering: a higher preference produces a higher probability. The denominator includes all available actions, so the probabilities are normalized across the action set for the state.

apply expcombine all actionsdivide by sumnormalizesPreferencesh(s, a₁, θ), h(s, a₂, θ),h(s, a₃, θ)Exponentialsexp(h) for each actionNormalization sumsum of all exp(h)Action probabilitiesπ(a | s, θ)
How do the numerical preferences for all available actions in one state become normalized action probabilities?

A Two-Action Softmax Calculation

For one state, suppose two actions have preferences h(s, a₁, θ) = ln(2) and h(s, a₂, θ) = 0.

Exponentiate each preference: The exponential preferences are exp(ln(2)) = 2 and exp(0) = 1.

Build the normalization sum: The denominator is 2 + 1 = 3.

Divide each exponential preference: Action a₁ receives 2/3 and action a₂ receives 1/3.

π(a₁ | s, θ) = 2/3 and π(a₂ | s, θ) = 1/3. The higher-preference action receives the higher probability, and the probabilities sum to 1.

Choosing a Preference Function

The softmax rule does not dictate how the preference values must be produced. The preference function h(s, a, θ) can be parameterized arbitrarily. A deep neural network is one possibility, with θ representing its vector of connection weights. A function that is linear in feature vectors φ(s, a) is another possibility. These choices change the computation of preferences, while the same softmax rule can convert those preferences into action probabilities.

represented asparameterized byprovided toproducesState-action pairs, aState-action pairs, aFeature vectorφ(s, a)Deep neural networkconnection weights θPreference h(s, a, θ)linear parameterizationPreference h(s, a, θ)neural parameterization
What is the difference between producing action preferences with a linear feature function and with a multilayer neural network?

The important design boundary is therefore clear: the parameterization determines how θ produces h(s, a, θ), and the softmax determines how h becomes a policy. Replacing a linear feature function with a deep neural network changes the preference-producing mechanism, not the basic preference-to-probability rule.

Why Differentiability Matters

Policy gradient methods adjust the policy through its weights. For that adjustment to be expressed as a gradient, the policy must be differentiable with respect to those weights. The required gradient must exist and must always be finite.

parameterizesdeterminesmust support differentiationWeights θpolicy parametersPreference h(s, a, θ)state-action preferenceSoftmax policyπ(a | s, θ)Policy gradientexists and is finite
How does a change in policy weights flow through the policy computation to change the probability of selecting an action?

Common Reasoning Errors

  • Treating a preference as an action probability

    Preferences are numerical inputs to the probability rule. They are not yet normalized action probabilities.

    Fix: Apply exponential softmax across all available actions before interpreting the results as probabilities.

  • Normalizing one action independently

    The softmax denominator uses the exponential preferences of all available actions for the state.

    Fix: Compute the shared normalization sum over the complete available-action set.

  • Assuming the preference function must be a neural network

    The preference function can be parameterized arbitrarily; a deep neural network and a linear feature function are both identified possibilities.

    Fix: Distinguish the mechanism that produces preferences from the softmax rule that converts them into probabilities.

  • Ignoring differentiability

    Policy gradient methods require a policy gradient that exists and is always finite.

    Fix: Check the differentiability requirement with respect to the policy weights.

  • Confusing preferences with action values

    A policy does not have to assign the same kind of value to actions that an action-value method does.

    Fix: Treat preferences as quantities used to rank actions and determine their selection probabilities.

Check Your Understanding

MEDIUM

A policy uses a preference function h(s, a, θ) for three actions in one state. Explain the complete path from policy weights θ to the final probability of selecting one action. Your explanation should identify the preference stage, the exponential softmax stage, and the differentiability requirement.

Hints
  • Start with how θ determines numerical preferences for state-action pairs.
  • State what the exponential softmax numerator and denominator contain.
  • Explain why the policy must have a gradient with respect to θ.

What do you think happens?

Suppose two actions have preferences ln(2) and 0. Before calculating, which action receives the larger probability, and what is the probability of each action after softmax?

  • The first action receives 2/3 and the second receives 1/3.
  • Both actions receive 1/2.
  • The first action receives 1/3 and the second receives 2/3.
Reveal answer

Answer: The first action receives 2/3 and the second receives 1/3.

Exponentiating the preferences gives 2 and 1. Their normalization sum is 3, so the probabilities are 2/3 and 1/3.

Key Takeaways

  • A parameterized policy uses weights θ to determine action probabilities.
  • The preference function h(s, a, θ) assigns numerical preferences to state-action pairs before probabilities are computed.
  • Exponential softmax gives higher probability to higher-preference actions and normalizes the probabilities across the available actions.
  • The preference function can be produced by a deep neural network, a linear function of feature vectors, or another parameterization.
  • Policy gradient methods require the policy to be differentiable with respect to its weights, with a gradient that exists and is always finite.