Average-Reward Reinforcement Learning
Differential values are defined with respect to the average reward.
The Continuing-Task Shift
Average-reward reinforcement learning addresses continuing problems: tasks that do not need to be organized around a terminal episode. The central change is how values are described. Instead of treating the episodic formulation as the only reference point, the continuing formulation defines differential values with respect to the average reward.
Reading Differential Values
The phrase differential values tells you how to interpret the value definitions in this setting. The values are defined with respect to the average reward. Average reward is therefore the reference named by the continuing formulation, and differential values are the corresponding value quantities under that definition.
This distinction is about definitions, not about abandoning value-based reasoning. When studying a continuing formulation, first ask which value definitions have changed. Only after that should you ask whether a theorem or equation that used those values also changes.
A Classification Walkthrough
Classifying what changes
Suppose a reinforcement-learning treatment moves from an episodic formulation to a continuing formulation. Classify the roles of the value definitions, the policy gradient theorem, and the forward and backward view equations.
Step 1: Identify the task setting: The continuing formulation does not need to be organized around a terminal episode.
Step 2: Identify the value change: The values are defined differentially with respect to the average reward.
Step 3: Check the policy gradient theorem: With the alternate value definitions, the policy gradient theorem remains true for continuing problems.
Step 4: Check the two views: The forward-view and backward-view equations remain the same in the continuing case.
The formulation changes its value definitions, but the cited theorem and the forward and backward view equations are preserved.
The Preserved Results
The policy gradient theorem from the episodic treatment remains true for the continuing case once the values have been given their continuing, differential definitions. This is the key bridge between the two settings: the task setting changes and the value definitions are adjusted, but the theorem still applies.
The same preservation applies to the forward-view and backward-view equations. In the continuing case, those equations remain the same. Therefore, do not infer that every equation must be rewritten merely because the value definitions are now differential.
Definition Versus Equation
| Question | Answer in the continuing formulation |
|---|---|
| What changes? | The values are defined differentially with respect to the average reward. |
| What happens to the policy gradient theorem? | It remains true for continuing problems. |
| What happens to the forward-view equations? | They remain the same. |
| What happens to the backward-view equations? | They remain the same. |
Common Misclassifications
Treating the continuing formulation as if it required terminal episodes.
A continuing problem does not need to be organized around a terminal episode.
Fix:
Begin by recognizing that the continuing setting uses differential value definitions with respect to the average reward.Assuming that differential values mean the entire learning framework has been replaced.
The policy gradient theorem remains true for continuing problems when the alternate value definitions are used.
Fix:
Separate the change in value definitions from the validity of the theorem.Rewriting the forward and backward view equations automatically.
The forward and backward view equations remain the same in the continuing case.
Fix:
Check each result individually instead of assuming that a changed definition changes every equation.
Check Your Understanding
A continuing reinforcement-learning problem is introduced. Decide which statement best describes the transition from the episodic treatment.
Hints
- Look first for the changed reference used to define values.
- Then check whether the policy gradient theorem is preserved.
- Finally, recall what happens to both the forward-view and backward-view equations.
What do you think happens?
Before revealing the answer, classify the following claim: In the continuing case, changing the value definitions necessarily changes the forward-view and backward-view equations.
Reveal answer
Answer: False
The values are defined differentially with respect to the average reward, but the forward-view and backward-view equations remain the same in the continuing case.
Key Takeaways
- Continuing problems do not need to be organized around terminal episodes.
- Their values are defined differentially with respect to the average reward.
- With these alternate value definitions, the policy gradient theorem remains true for continuing problems.
- The forward-view and backward-view equations remain the same in the continuing case.
- The safest analysis separates changes in value definitions from results and equations that remain preserved.
Key Takeaways
- Differential values use the average reward as their reference in continuing problems.
- The continuing formulation changes how values are defined, not the entire reinforcement-learning framework.
- The policy gradient theorem remains valid in the continuing case.
- The forward-view and backward-view equations remain the same.
- Always distinguish a definition change from a change in the equations that use the definition.