Generalization and Function Approximation in Reinforcement Learning
A policy can be represented by a parameterized function whose parameters form the weight vector θ.
The Scale Problem
A reinforcement-learning agent may need to act in an environment with far more possible states than it can practically list, store, and process one by one. A complete and accurate model would not automatically solve this problem: the agent might still lack enough memory to store all the relevant information or enough time to perform the necessary computations at every step. Function approximation addresses this scale problem by representing information with a parameterized function instead of a separate exact entry for every possible state.
From State to Policy
A policy can be represented by a parameterized function. The parameters of that function form the weight vector, written as θ. Read the representation as a chain: a state is supplied to the parameterized function, θ identifies the particular parameter setting, and the resulting representation determines the policy's action behavior. Depending on the policy representation, the result may be action probabilities or a selected action.
Reading a Policy Representation
Suppose two candidate weight vectors, θA and θB, are used with the same parameterized policy function.
Identify the state: The function receives a state as the situation in which the agent must choose how to behave.
Identify the parameter setting: θA and θB specify two different parameter settings for the policy representation.
Compare the resulting behavior: Each parameter setting can produce its own action probabilities or selected action for that state.
The weight vector is not an isolated answer for one state. It identifies a parameterized policy representation whose behavior is evaluated across states.
Why the Representation Generalizes
The weight-vector representation is approximate because the state space is much larger than the weight vector. The agent is not storing a separate exact answer for every state. Instead, one parameter setting controls a function that represents the policy and associated values across the state space. As a result, changing the shared weight vector can alter the estimated values or action preferences for multiple states rather than changing only one separately stored entry. This is the central generalization idea: the compact representation is used across many states, but it cannot be interpreted as a complete list of exact state-by-state information.
Reading MSVE
MSVE evaluates the error in the values associated with a chosen weight vector θ under the on-policy distribution d. In the source notation, vπθ(s) denotes the values associated with the policy represented by θ, while d supplies the state weighting used when measuring the error.
To interpret MSVE, keep three objects separate. First, θ is the weight vector. Second, vπθ(s) denotes the value approximation associated with the policy represented by that vector. Third, d is the on-policy distribution. MSVE combines the value errors across states while weighting those errors according to d. Therefore, it is not merely a visual comparison of two weight vectors and not a count of how many parameters they contain. It is an evaluation of the values induced by each parameter setting under the distribution generated by the policy.
Comparing Candidate Weights
An On-Policy Comparison
An agent evaluates two candidate weight vectors, θA and θB, for the same value-function approximation setting and the same on-policy distribution d.
Evaluate θA: Use θA to produce the values associated with its represented policy. Compare those values with the corresponding true values across the states considered by d.
Evaluate θB: Repeat the same process for θB. Its values must be evaluated over the same on-policy basis so that the comparison is meaningful.
Use MSVE: MSVE combines the state-by-state squared value errors while weighting states according to d, producing one error measure for θA and another for θB.
Rank the approximations: The candidate with the lower MSVE has the lower measured value error under that on-policy distribution.
MSVE supplies a common on-policy measure for comparing the value-function approximations induced by θA and θB. The comparison is about their induced values and errors, not about which weight vector looks shorter or simpler.
Optimality Under Constraints
Memory and computation limits affect more than the storage of state information. They can also limit the agent's ability to represent value functions, policies, and models, and can prevent the agent from fully using even a complete and accurate environment model. For this reason, an agent may work with approximations of all three. Optimality remains a useful theoretical target: it gives a reference for what the agent is trying to approach. In practice, the agent may achieve only an approximation of the optimal value or policy because the available representation, memory, or computation is limited.
When analyzing an approximation, state both the target and the evaluation setting. Identify what optimality would mean as a theoretical ideal, identify the parameterized representation actually available, and use an appropriate measure such as on-policy MSVE to assess how closely the represented values match the target in the relevant distribution.
Mistakes in Reasoning
Treating the weight vector as a table containing one exact answer for every state.
The weight vector is a compact parameter setting for a function, while the state space is much larger than the weight vector.
Fix:
Describe θ as controlling a parameterized policy or value representation that produces approximations across states.Comparing two weight vectors by their appearance or length.
MSVE compares the errors in the associated values under the on-policy distribution, not the visual form of the parameter vectors.
Fix:
Evaluate the values associated with each candidate and compare their on-policy MSVE scores.Ignoring the distribution used by MSVE.
The source defines MSVE using the on-policy distribution d, which weights the state errors.
Fix:
State that the comparison is made under d and interpret the score in that on-policy setting.Assuming that a complete model removes all practical limitations.
Limited computation can prevent an agent from fully using even a complete and accurate model.
Fix:
Consider both what the agent knows and what it can store and compute.Treating optimality as something a practical agent must always achieve exactly.
Optimality is a useful theoretical target, while real agents generally achieve only an approximation because of resource and representation limits.
Fix:
Distinguish the ideal target from the quality of the approximation that can actually be represented and computed.
Check Your Understanding
An agent has two candidate weight vectors, θ1 and θ2. Both are used with the same parameterized value representation and evaluated under the same on-policy distribution d. Explain the sequence of reasoning you would use to decide which candidate is better according to MSVE. Then explain why the result does not prove that the chosen candidate is exact for every possible state.
Hints
- Begin with the values associated with the policy represented by each weight vector.
- Compare each candidate's value errors with the corresponding true values.
- Remember that d determines how the state errors are weighted.
- A lower on-policy MSVE indicates lower measured error in that evaluation setting, not exact storage of every state value.
What do you think happens?
Two candidate weight vectors are evaluated with the same on-policy distribution. Candidate θA has lower MSVE than candidate θB. Which candidate has the lower measured value error in that on-policy comparison?
Reveal answer
Answer: θA
MSVE evaluates the value errors associated with a chosen weight vector under the on-policy distribution. A lower MSVE therefore indicates lower measured error in that comparison.
Key Takeaways
- A policy can be represented by a parameterized function whose parameters form the weight vector θ.
- The representation is approximate because the state space can be much larger than the weight vector, so it does not store a separate exact answer for every state.
- MSVE evaluates the value errors associated with θ while weighting states according to the on-policy distribution d.
- MSVE provides a common basis for comparing value-function approximations in the on-policy case.
- Memory and computation limits affect the representation of states, value functions, policies, and models; optimality is therefore a useful target even when only an approximation can be achieved.
Key Takeaways
- A parameterized function uses a weight vector θ to represent a policy and its associated values across states.
- Because the state space may be much larger than the weight vector, the representation is an approximation rather than a separate exact state table.
- MSVE measures the state-weighted value error for a chosen θ under the on-policy distribution d.
- Comparing MSVE scores allows candidate value-function approximations to be ranked in the on-policy setting.
- Optimality is a theoretical reference point; practical agents often use approximate values, policies, and models because memory and computation are limited.