On-Policy Prediction with Function Approximation
Function approximation introduces distinctive challenges into reinforcement learning.
Why the Setting Changes
In reinforcement learning, prediction means learning values from experience. Function approximation changes the setting because predictions are no longer treated only as isolated entries in a table. The important lesson is not that one difficulty replaces another: nonstationarity, bootstrapping, and delayed targets are separate issues that must be examined individually. These issues make reinforcement learning with function approximation different from conventional supervised learning.
Shared Representation
A useful conceptual starting point is that function approximation represents predictions through a general approximation mechanism rather than treating every prediction as an unrelated item. In that setting, a change made while learning one prediction can be relevant to other predictions represented by the same approximation mechanism. This is the intuition behind the challenge: learning is no longer just a matter of correcting one isolated value. The source pack identifies this broader setting as one in which distinctive reinforcement-learning issues arise.
A conceptual shared-representation example
Suppose several states are represented through one learned approximation mechanism. Why should a learner avoid thinking of each prediction as an entirely isolated table entry?
Identify the representation: The predictions are produced through function approximation rather than being treated only as separate tabular entries.
Notice the shared mechanism: The same approximation mechanism is involved in representing predictions for multiple states.
Reconsider the update: A learning change should be understood as a change to the approximation mechanism, not automatically as a correction confined to one independent entry.
Connect the idea to the chapter: This broader representation is why the learning objectives and the usual assumptions must be examined again.
Function approximation changes the structure of the prediction problem: predictions are connected through a shared approximation mechanism, so the learning analysis must account for that setting.
Three Separate Difficulties
The source pack emphasizes three issues that should not be collapsed into one: nonstationarity, bootstrapping, and delayed targets. Treating them separately is important because each names a different reason that learning can be difficult. The source does not present them as interchangeable labels. Instead, they are separate issues within the broader challenge of reinforcement learning with function approximation.
Treating nonstationarity, bootstrapping, and delayed targets as one single problem.
The source identifies the three as separate issues.
Fix:
Name and examine each issue independently before discussing the learning objective.Assuming that objectives from an earlier learning setting automatically remain appropriate.
The learning objectives must be reconsidered in this setting.
Fix:
Treat the objective itself as part of the analysis.Assuming reinforcement learning with function approximation has only the challenges found in conventional supervised learning.
The source states that reinforcement learning with function approximation introduces issues that do not normally arise in conventional supervised learning.
Fix:
Analyze reinforcement learning with function approximation as its own setting.
On-Policy Update Loop
The chapter begins with on-policy prediction and control. At a high level, the prediction system follows the policy being considered, receives experience, forms learning information, and updates its value-function approximation. The key educational point is the loop itself: experience and approximation interact repeatedly, so the learner must analyze both the source of the experience and the way predictions are used during learning.
This loop should be read as a conceptual trace, not as a complete algorithm specification. The source pack identifies bootstrapping and delayed targets as separate issues, but it does not provide a particular update rule. Therefore, the important conclusion here is analytical: when predictions and targets interact over successive updates, the learner must reconsider the objectives rather than assume that the tabular analysis transfers unchanged.
Practice and Transfer
A study group says: Function approximation is merely a compact way to store values, so the same learning objectives can be used without further analysis. Evaluate this statement using the three separate issues named in the source pack.
Hints
- Identify the setting being discussed.
- List the three issues separately rather than treating them as one.
- Explain why the learning objectives must be reconsidered.
Evaluating the study-group claim
Decide whether the statement that function approximation leaves the learning objectives unchanged is justified.
Check the setting: The statement concerns reinforcement learning with function approximation, not only conventional supervised learning.
Separate the issues: The relevant issues are nonstationarity, bootstrapping, and delayed targets. They should be treated as separate rather than hidden under one general label.
Check the objective: The source explicitly states that learning objectives must be reconsidered in this setting.
Reach the conclusion: The study group's claim is too strong because it assumes unchanged objectives where the source requires renewed examination.
The claim is not justified. Function approximation introduces distinctive reinforcement-learning challenges, and the learning objectives must be reconsidered.
When analyzing an on-policy prediction method with function approximation, document four things separately: the policy whose experience is being considered, the approximation mechanism used for predictions, which of the three named issues is present, and whether the learning objective has been explicitly reconsidered.
Learning Path Ahead
On-policy prediction and control form the starting point for the broader treatment. The topic then progresses to off-policy methods, eligibility traces, and policy-gradient methods. Understanding the initial distinction between function approximation and conventional supervised learning prepares you to examine those later settings without assuming that their objectives and difficulties are identical.
- Function approximation introduces distinctive challenges into reinforcement learning.
- Nonstationarity, bootstrapping, and delayed targets are separate issues and should be analyzed separately.
- Reinforcement learning with function approximation raises issues that do not normally arise in conventional supervised learning.
- Learning objectives must be reconsidered rather than assumed to transfer unchanged.
- The treatment begins with on-policy prediction and control before moving to off-policy methods, eligibility traces, and policy-gradient methods.
Key Takeaways
- Function approximation changes the analysis of reinforcement-learning prediction.
- Nonstationarity, bootstrapping, and delayed targets are distinct issues.
- On-policy prediction provides the starting point for examining these challenges.
- The learning objectives must be reconsidered in the function-approximation setting.
- Later topics extend the treatment to off-policy methods, eligibility traces, and policy-gradient methods.