Concepts / Parameterized Function Approximation

Parameterized Function Approximation

Continuing cases motivate the average reward formulation.

  • Programming

Why Continuing Control Changes the Objective

Some control problems do not naturally end. The interaction continues, so the learning objective must describe performance across an ongoing sequence of time steps rather than only within an episode. This motivates the average reward formulation: instead of relying on an episodic endpoint, control is framed around maximizing the average reward per time step.

continuescontinuescontributes toTime step 1rewardTime step 2rewardTime step 3rewardAverage reward pertime stepcontinuing objective
How does reward accumulate over an ongoing interaction, and how is the long-run average reward per time step used?

For episodic control, parameterized function approximation and semi-gradient descent extend naturally. Continuing control is different because it requires a problem formulation based on maximizing average reward per time step.

The Approximate-Control Limitation

The discounted formulation cannot simply be carried over to control when approximations are present. In the exact case, control can use value functions to guide comparisons. In the approximate case, most policies cannot be represented by a value function in the required way. That creates a basic control problem: policies that lack such a representation still need to be compared and ranked.

supportsrequires another basisExact controlvalue-functionrepresentationApproximate controlmost policies notrepresentedPolicy comparisonstill required
What assumption makes discounted control work in the exact case, and what changes when value functions are represented approximately?

Ranking Policies with Average Reward

The average reward formulation supplies a scalar quantity, written as η(π), for a policy π. Because η(π) is a single average-reward measure, it can be used to rank arbitrary policies. This separates two tasks that are easy to confuse: representing the values associated with a policy and deciding which policy is better.

is represented byhassupportsPolicy πValue representationvalues associated with πPolicy rankingcompare policiesη(π)scalar average reward
How can approximate value estimates represent a policy while the scalar average reward η(π) independently ranks different policies?

Using η(π) to Compare Policies

Suppose two arbitrary policies, πA and πB, do not both have usable value-function representations in the approximate setting. How can they still be ordered?

Identify the limitation: A value-function representation is not available for most policies in the approximate case, so it cannot serve as the common comparison basis.

Assign the scalar measure: Associate η(πA) with πA and η(πB) with πB. Each quantity is the policy's average reward per time step.

Rank the policies: Compare the two scalar average rewards. The policy with the higher η value is ranked higher by the average reward criterion.

η(π) provides a common scalar basis for ranking arbitrary policies, even when value-function representation is unavailable for most of them.

The source example does not require numerical reward data. The important idea is that η(π) changes the task from requiring a value-function representation for most policies to ranking policies with one scalar quantity.

From States to Shared Weights

A reinforcement-learning policy may need to account for a state space far larger than any practical list of individual states. Parameterized function approximation addresses this scale problem by representing the policy with a function controlled by a weight vector, written as θ. The state is handled through the parameterized function, while θ identifies the particular parameter setting being considered.

inputcontrolsproducesStatesParameterizedfunctionuses s and θPolicy or valueestimateassociated with θWeight vectorθ
How does a parameterized function use a state as input and a weight vector to produce a policy or value estimate?
could requiredeterminesLarge state spacemany possible statesSeparate stateentriesone per stateWeight vector θcompact representationApproximate valuesnot one exact answer perstate
What information is retained or lost when many state values are represented by a smaller set of shared weights instead of one value per state?

Measuring Approximation with MSVE

For a chosen weight vector θ, MSVE evaluates the error in the values associated with the policy represented by θ, weighting states according to the on-policy distribution d. It provides a performance measure for the value approximation rather than leaving the choice of θ to intuition.

Keep three objects separate. θ is the weight vector. vπθ(s) denotes the values associated with the policy represented by θ. The distribution d is the on-policy distribution used when measuring error. MSVE connects these pieces by evaluating how far the approximate values induced by θ are from the relevant true values under d.

inducescompared with true valuesreferenceweights statesθweight vectorvπθ(s)approximate valuesMSVE(θ)weighted value errorTrue valuespolicy valuesdon-policy distribution
How does MSVE measure the weighted difference between an approximate value function and the true value function under the policy's state distribution?

Comparing Two Weight Vectors

Two candidate weight vectors, θA and θB, define two parameterized value approximations for the same on-policy setting. What is the correct comparison?

Generate the associated values: Treat each weight vector as defining its own parameterized policy representation and associated values.

Use the same evaluation idea: Evaluate the approximation induced by θA and the approximation induced by θB using MSVE under the on-policy distribution d.

Compare the errors: The weight vector with the lower MSVE gives the better value-function approximation according to this on-policy measure.

MSVE compares the value approximations induced by different weight vectors; the vectors themselves should not be judged merely by their lengths or appearances.

Mistakes About Approximate Control

  • Assuming the discounted formulation transfers directly to approximate control.

    The source identifies the lack of a value-function representation for most policies as the limitation that blocks this direct carry-over.

    Fix: Use the average reward formulation for continuing control and use η(π) to rank arbitrary policies.

  • Treating value representation and policy ranking as the same task.

    The scalar η(π) provides a separate basis for ranking policies.

    Fix: Keep the value representation and the scalar policy-ranking measure conceptually distinct.

  • Interpreting θ as a separate exact value for every state.

    The state space is much larger than the weight vector, so the representation is an approximation.

    Fix: Treat θ as the compact parameter setting used by a function to produce associated policy or value estimates.

  • Comparing weight vectors by their appearance instead of their induced values.

    MSVE evaluates the values associated with each parameter vector under the on-policy distribution.

    Fix: Compare the induced value approximations using MSVE.

When applying this method, name the objects explicitly: θ is the weight vector, vπθ(s) is the value associated with the policy represented by θ, and d is the on-policy distribution used by MSVE. This prevents the representation, the policy, and the evaluation measure from being merged into one vague idea.

Check Your Understanding

MEDIUM

A continuing control problem uses parameterized function approximation. Explain why a scalar η(π) is useful even when most policies cannot be represented by a value function. Then describe how MSVE would compare two candidate weight vectors in the on-policy case.

Hints
  • First separate the problem of representing a policy from the problem of ranking policies.
  • Remember that η(π) is the policy's scalar average reward measure.
  • For MSVE, identify θ, the associated values vπθ(s), and the on-policy distribution d.

What do you think happens?

Two candidate weight vectors represent two value-function approximations for the same on-policy setting. If one has lower MSVE than the other, which approximation is preferred by this measure?

  • The approximation with lower MSVE
  • The approximation with the longer weight vector
  • The approximation whose weights look simpler
Reveal answer

Answer: The approximation with lower MSVE

MSVE measures the weighted value error under the on-policy distribution, so the lower-error approximation is preferred by this evaluation measure.

Key Takeaways

  • Continuing control motivates an objective based on maximizing average reward per time step.
  • The discounted formulation does not carry over directly to approximate control because most policies cannot be represented by a value function.
  • The scalar η(π) provides a common way to rank arbitrary policies even when value-function representation is unavailable.
  • A parameterized function uses a weight vector θ to represent policies or associated values compactly when the state space is much larger than the vector.
  • MSVE evaluates the weighted value error induced by θ under the on-policy distribution d, allowing value-function approximations to be compared.