Concepts / Discounted Formulation

Discounted Formulation

Continuing cases motivate the average reward formulation.

  • Programming

The Continuing-Control Problem

Some control problems are episodic: they have a natural episode structure, and parameterized function approximation together with semi-gradient descent extends naturally to control in that setting. Continuing problems are different. They do not use the same episodic framing, so they require a new problem formulation based on maximizing the average reward per time step.

continues acrosssummarized byReward streamcontinuing taskTime stepsongoingAverage reward pertime stepobjective
How does a continuing stream of rewards become one long-run quantity when there is no terminal episode?

The important shift is the objective. In continuing control, the task is formulated around the reward obtained on average at each time step. This gives the learner a way to evaluate ongoing behavior rather than relying on an episodic control formulation.

Two Formulations for Control

FormulationSetting described by the sourceMain role
Discounted formulationEpisodic controlSupports the episodic extension of parameterized function approximation and semi-gradient descent
Average reward formulationContinuing controlMakes maximizing average reward per time step the control objective
supportsmaximizesDiscountedformulationepisodic controlAverage rewardformulationcontinuing controlFunctionapproximation andsemi-gradient descentextends naturallyAverage reward pertime stepcontrol objective
How do the two formulations evaluate control in the settings described by the source?

Where Approximation Breaks the Argument

The discounted formulation has a limitation when control uses approximation. In the approximate case, most policies cannot be represented by a value function. That creates a direct problem: control still needs to compare policies, but a value-function representation is unavailable for most of the policies being considered.

extendsencountersleavesEpisodic controlapproximation extendsnaturallyApproximate controldiscounted formulationFunctionapproximationwith semi-gradient descentMost policiesnot represented by a valuefunctionPolicy comparisonstill required
What part of the discounted-control argument breaks when value functions and policies are only approximated?

Approximate-control limitation: the discounted formulation cannot be carried over to control when approximations are present because most policies cannot be represented by a value function, even though those policies still need to be compared and ranked.

Ranking Policies with η(π)

Replacing Representation with a Common Ranking Quantity

Suppose an approximate-control problem considers several arbitrary policies. The value-function representation is unavailable for most of them. What role can η(π) play?

Identify the limitation: Most policies cannot be represented by a value function in the approximate case, so value-function representation cannot serve as the common basis for comparing all policies.

Assign the scalar quantity: The average reward formulation associates each policy π with a scalar average reward, written as η(π).

Use a common comparison: The scalar η(π) gives a common quantity for ordering arbitrary policies, including policies that lack the required value-function representation.

Select by ranking: Policies can be ranked according to their η(π) values. The example does not calculate numerical values; it illustrates the role of the scalar quantity.

η(π) changes the control question from requiring a value-function representation for most policies to ranking arbitrary policies with a scalar average reward.

mapped tomapped tomapped toPolicy π1arbitrary policyη(π1)average rewardPolicy π2arbitrary policyη(π2)average rewardPolicy π3arbitrary policyη(π3)average reward
How can average reward values for several arbitrary policies be compared to determine which policy is better?

The key point is not that η(π) describes every internal detail of a policy. Its role is to provide one scalar average reward for that policy. Those scalar quantities can then be used to order arbitrary policies, even when most of the policies cannot be represented by value functions.

Representation and Ranking

may be represented byis assignedsupportsPolicy πbehavior to evaluateValue functionrepresentationPolicy rankingcomparison across policiesη(π)scalar average reward
How can an approximate value function represent a policy's behavior while η(π) separately determines how policies are ranked?

Value-function representation and policy ranking are related but different tasks. A value function is a representation issue: can the policy be represented in that form? η(π) addresses the ranking issue: once policies are associated with scalar average rewards, how can they be ordered? In the approximate case, the first task may fail for most policies, while the second task still has to be solved.

Common Misunderstandings

  • Treating continuing control as if it were simply the episodic case extended indefinitely.

    The source states that continuing cases require a new problem formulation based on maximizing average reward per time step.

    Fix: Start by identifying the setting. For continuing control, use the average reward formulation as the objective described by the source.

  • Assuming the discounted formulation can always be carried over when approximation is used.

    In the approximate case, most policies cannot be represented by a value function.

    Fix: Recognize the limitation and use η(π) as a scalar quantity for ranking arbitrary policies.

  • Treating η(π) as another name for a value-function representation.

    η(π) is a scalar average reward used to rank policies; the source distinguishes this role from representing a policy with a value function.

    Fix: Separate representation from ranking. A policy may lack the relevant value-function representation while still being assigned an η(π) for comparison.

  • Expecting the concept example to require numerical reward data.

    The source provides no reward sequence or numerical policy data and uses the example only to show the role of η(π).

    Fix: Focus on the structural idea: η(π) supplies a common scalar with which arbitrary policies can be ordered.

Apply the Distinction

MEDIUM

Explain, in your own words, why a continuing approximate-control problem needs both a way to handle policy representation and a separate way to rank policies. Your answer should mention the limitation of value-function representation and the role of η(π).

Hints
  • Begin with what happens to most policies in the approximate case.
  • Then explain why control still needs to compare those policies.
  • Finish by describing η(π) as a scalar average reward used for ranking.

What do you think happens?

A continuing approximate-control problem contains policies that cannot be represented by value functions. Can those policies still be ranked using the average reward formulation?

  • Yes, using η(π)
  • No, because ranking requires a value-function representation
Reveal answer

Answer: Yes, using η(π)

The source identifies η(π) as a scalar average reward that provides an effective way to rank arbitrary policies, including policies that cannot be represented by a value function.

Key Takeaways

  1. Continuing control requires a formulation based on maximizing average reward per time step.
  2. The discounted formulation extends naturally with parameterized function approximation and semi-gradient descent in the episodic case, but it cannot be carried over to approximate control in the continuing case.
  3. In approximate control, most policies cannot be represented by a value function.
  4. The scalar η(π) supplies a common average-reward quantity for ranking arbitrary policies.
  5. Value-function representation and policy ranking are separate roles: the first may be unavailable for most policies, while η(π) still supports comparison.

Key Takeaways

  • Continuing tasks motivate an average reward formulation because the objective is average reward per time step.
  • Approximation creates a limitation: most policies cannot be represented by a value function.
  • Control still needs to compare those policies, so representation alone is not enough.
  • The scalar η(π) provides a common quantity for ranking arbitrary policies.
  • The average reward formulation separates the question of representing a policy from the question of ranking it.