Concepts / Value Functions

Value Functions

A policy gradient method searches over parameter-defined policies rather than treating a value estimate as the primary search object.

  • Machine Learning

From Values to Policies

Many reinforcement learning discussions begin by asking how good a situation or action is. Policy gradient methods begin from a different angle: they search directly through possible policies. Each policy is described by a collection of numerical parameters. The method uses the agent's interaction with its environment to estimate how those parameters should change if the policy is to perform better.

definescan be adjusted to definecan be adjusted to definePolicy parametersnumerical settingsPolicy Aone parameter settingPolicy Banother parameter settingPolicy Canother parameter setting
How do numerical parameters define the policies being searched?

The primary search object is a policy described by numerical parameters, not a value estimate. The central estimate in a policy gradient method is a direction for changing those parameters.

The Policy-Environment Loop

To trace a policy gradient method, follow information from the policy into the environment and back again. Numerical parameters first define the policy being used. The agent then interacts with the environment. Those behavioral interactions provide information that can be used to estimate a direction for improving the policy. The parameters are adjusted in that direction, producing another parameter-defined policy, and the search can repeat.

defineguides interactionprovides information forguides adjustment ofdefinePolicy parameterscurrent valuesPolicydefined by parametersEnvironmentagent interactionImprovement directionestimated from interactionAdjusted parametersnew valuesNew policyanother parameter-definedpolicy
How does experience from the environment lead to a new policy?

The environment interaction matters because the improvement direction is not described as an arbitrary change to the parameters. It is estimated from the agent's behavioral interaction with the environment. Without that interaction, the method would not have the experience-based information described in the source.

What Gets Adjusted

Tracing One Policy Update

Suppose an agent currently uses one policy defined by numerical parameters. Trace what changes when the method estimates that a different parameter setting would improve performance.

Start with parameters: The agent has numerical parameters that define the policy it currently follows.

Interact with the environment: The agent uses that policy while interacting with the environment. The resulting behavioral interactions provide information about how the policy is performing.

Estimate a direction: The method estimates a direction for changing the numerical parameters.

Adjust the parameters: The parameters are moved in the estimated direction. The meaningful change is in the numerical parameters, not a switch between two named algorithms.

Define another policy: After adjustment, the parameters define another policy. The method can repeat this search process.

A policy gradient update changes the numerical parameters that define the policy, using information from interaction with the environment.

definedefineParameterscurrent settingParametersadjusted settingPolicydefined by current settingPolicydefined by adjusted setting
What is the meaningful state change during a policy gradient search?

Value Estimates and Gradients

Value function estimates can appear inside some policy gradient methods, but they are not required by every method in the family. Their role is supportive: they can improve the estimated gradient, meaning they can make the direction for changing policy parameters more useful. A value estimate and a policy gradient estimate are therefore related but not identical.

can informinformscan improveguidesEnvironmentinteractionexperienceValue estimateused by some methodsGradient estimatedirection for parameterchangeParameter adjustmentsearch continues
Where can a value-function estimate enter the policy-gradient process?
  • Treating the value function as the primary search object in every policy gradient method.

    The source identifies the policy, not the value estimate, as the object being searched through. It also states that value functions are not required by every method in this family.

    Fix: Say that the method searches through parameter-defined policies. Add that some methods use value estimates to improve gradient estimates.

  • Equating a value estimate with the policy gradient estimate.

    The value estimate can improve the gradient estimate, but the gradient estimate is the direction for changing the policy parameters.

    Fix: Keep the roles separate: the value estimate may support the gradient estimate, while the gradient estimate supplies the parameter-change direction.

  • Ignoring environment interaction.

    The source states that the direction is estimated from the agent's interaction with the environment.

    Fix: Include the flow from policy to environment interaction, then from interaction to an estimated improvement direction.

State Value and Action Value

In a Markov decision process, value functions help describe how valuable a situation or decision is under a policy. A state-value function evaluates a state s under a policy π. It asks about the value of being in that state when the policy is the policy being considered.

The action-value function qπ(s, a) evaluates a more specific situation. It represents the expected return from starting at state s, taking action a, and thereafter following policy π. Its description fixes three parts: the starting state, the action taken first, and the policy followed afterward.

evaluatesunderstarts attakes firstfollows afterwardState valuestate s under policy πAction valueqπ(s, a)State sincludedState sstarting statePolicy πincludedAction aselected firstPolicy πfollowed afterward
What changes when the evaluation includes one specific initial action?
FunctionWhat is fixed or specified?What is evaluated?
State-value functionState s and policy πThe state under that policy
Action-value function qπ(s, a)Starting state s, first action a, and later policy πThe expected return for that specific initial action followed by the policy

The essential difference is whether a specific initial action is included.

Reading qπ(s, a)

Unpacking the Action-Value Description

Interpret qπ(s, a) without losing track of what is fixed.

Identify s: The evaluation starts from the specified state s.

Identify a: The agent takes the specified action a first. This is the detail that makes action value more specific than state value.

Identify π: After that initial action, the agent follows policy π.

Read the result: The function evaluates the expected return for that complete description: start at s, take a, and thereafter follow π.

qπ(s, a) is the expected return starting from state s, taking action a, and thereafter following policy π.

takethen followevaluateState sstarting stateAction aselected firstPolicy πfollowed afterwardExpected returnqπ(s, a)
What is fixed when evaluating qπ(s, a), and how does the selected action lead into following policy π afterward?

Practice Check

MEDIUM

Explain the difference between a state-value function and qπ(s, a) in two or three sentences. Your answer should identify the starting state, the initial action when one is specified, and the policy followed afterward.

Hints
  • Begin with what a state-value function evaluates.
  • Then explain what extra detail appears in qπ(s, a).
  • Use the sequence: start at s, take a, then follow π.
MEDIUM

A policy gradient method has a current policy defined by numerical parameters. Describe the information flow that can lead to a new policy. Include the environment interaction and explain where a value estimate may help.

Hints
  • Start with parameters defining the current policy.
  • Include interaction with the environment.
  • Separate the gradient estimate from an optional value-function estimate.

Key Takeaways

  1. Policy gradient methods search through policies defined by numerical parameters.
  2. Interaction with the environment provides information used to estimate a direction for changing those parameters.
  3. Some methods use value-function estimates to improve gradient estimates, but value functions are not required by every method in the family.
  4. A state-value function evaluates a state under a policy.
  5. qπ(s, a) evaluates the expected return from starting at s, taking a first, and thereafter following π.

Key Takeaways

  • Policy gradient methods search over parameter-defined policies rather than treating value estimates as the primary search object.
  • The estimated improvement direction comes from interaction with the environment and is used to adjust policy parameters.
  • Value estimates can support and improve gradient estimates in some methods, but they are not universal requirements.
  • State value evaluates a state under a policy, while qπ(s, a) evaluates a specified initial action from that state followed by the policy.
  • The action-value description fixes the starting state, the selected first action, and the policy followed afterward.