Discounted Formulation
Continuing cases motivate the average reward formulation.
The Continuing-Control Problem
Some control problems are episodic: they have a natural episode structure, and parameterized function approximation together with semi-gradient descent extends naturally to control in that setting. Continuing problems are different. They do not use the same episodic framing, so they require a new problem formulation based on maximizing the average reward per time step.
The important shift is the objective. In continuing control, the task is formulated around the reward obtained on average at each time step. This gives the learner a way to evaluate ongoing behavior rather than relying on an episodic control formulation.
Two Formulations for Control
| Formulation | Setting described by the source | Main role |
|---|---|---|
| Discounted formulation | Episodic control | Supports the episodic extension of parameterized function approximation and semi-gradient descent |
| Average reward formulation | Continuing control | Makes maximizing average reward per time step the control objective |
Where Approximation Breaks the Argument
The discounted formulation has a limitation when control uses approximation. In the approximate case, most policies cannot be represented by a value function. That creates a direct problem: control still needs to compare policies, but a value-function representation is unavailable for most of the policies being considered.
Approximate-control limitation: the discounted formulation cannot be carried over to control when approximations are present because most policies cannot be represented by a value function, even though those policies still need to be compared and ranked.
Ranking Policies with η(π)
Replacing Representation with a Common Ranking Quantity
Suppose an approximate-control problem considers several arbitrary policies. The value-function representation is unavailable for most of them. What role can η(π) play?
Identify the limitation: Most policies cannot be represented by a value function in the approximate case, so value-function representation cannot serve as the common basis for comparing all policies.
Assign the scalar quantity: The average reward formulation associates each policy π with a scalar average reward, written as η(π).
Use a common comparison: The scalar η(π) gives a common quantity for ordering arbitrary policies, including policies that lack the required value-function representation.
Select by ranking: Policies can be ranked according to their η(π) values. The example does not calculate numerical values; it illustrates the role of the scalar quantity.
η(π) changes the control question from requiring a value-function representation for most policies to ranking arbitrary policies with a scalar average reward.
The key point is not that η(π) describes every internal detail of a policy. Its role is to provide one scalar average reward for that policy. Those scalar quantities can then be used to order arbitrary policies, even when most of the policies cannot be represented by value functions.
Representation and Ranking
Value-function representation and policy ranking are related but different tasks. A value function is a representation issue: can the policy be represented in that form? η(π) addresses the ranking issue: once policies are associated with scalar average rewards, how can they be ordered? In the approximate case, the first task may fail for most policies, while the second task still has to be solved.
Common Misunderstandings
Treating continuing control as if it were simply the episodic case extended indefinitely.
The source states that continuing cases require a new problem formulation based on maximizing average reward per time step.
Fix:
Start by identifying the setting. For continuing control, use the average reward formulation as the objective described by the source.Assuming the discounted formulation can always be carried over when approximation is used.
In the approximate case, most policies cannot be represented by a value function.
Fix:
Recognize the limitation and use η(π) as a scalar quantity for ranking arbitrary policies.Treating η(π) as another name for a value-function representation.
η(π) is a scalar average reward used to rank policies; the source distinguishes this role from representing a policy with a value function.
Fix:
Separate representation from ranking. A policy may lack the relevant value-function representation while still being assigned an η(π) for comparison.Expecting the concept example to require numerical reward data.
The source provides no reward sequence or numerical policy data and uses the example only to show the role of η(π).
Fix:
Focus on the structural idea: η(π) supplies a common scalar with which arbitrary policies can be ordered.
Apply the Distinction
Explain, in your own words, why a continuing approximate-control problem needs both a way to handle policy representation and a separate way to rank policies. Your answer should mention the limitation of value-function representation and the role of η(π).
Hints
- Begin with what happens to most policies in the approximate case.
- Then explain why control still needs to compare those policies.
- Finish by describing η(π) as a scalar average reward used for ranking.
What do you think happens?
A continuing approximate-control problem contains policies that cannot be represented by value functions. Can those policies still be ranked using the average reward formulation?
Reveal answer
Answer: Yes, using η(π)
The source identifies η(π) as a scalar average reward that provides an effective way to rank arbitrary policies, including policies that cannot be represented by a value function.
Key Takeaways
- Continuing control requires a formulation based on maximizing average reward per time step.
- The discounted formulation extends naturally with parameterized function approximation and semi-gradient descent in the episodic case, but it cannot be carried over to approximate control in the continuing case.
- In approximate control, most policies cannot be represented by a value function.
- The scalar η(π) supplies a common average-reward quantity for ranking arbitrary policies.
- Value-function representation and policy ranking are separate roles: the first may be unavailable for most policies, while η(π) still supports comparison.
Key Takeaways
- Continuing tasks motivate an average reward formulation because the objective is average reward per time step.
- Approximation creates a limitation: most policies cannot be represented by a value function.
- Control still needs to compare those policies, so representation alone is not enough.
- The scalar η(π) provides a common quantity for ranking arbitrary policies.
- The average reward formulation separates the question of representing a policy from the question of ranking it.