Stochastic and Semi-gradient Methods for Value Prediction
SGD applies a sequence of parameter updates to function approximation in value prediction.
From State Values to Adjustable Parameters
Suppose an agent wants to predict the value of a state under a policy. Each state may have a correct value, but the setting considered here does not store a separate perfect answer for every possible state. Instead, the agent uses a function approximator controlled by a finite collection of real-valued parameters. Learning consists of adjusting those parameters so that the approximator produces useful value predictions.
The approximate value function is written v̂(s, θ). It receives a state s and the current weight vector θ, then produces an estimated value. The weight vector is the adjustable part of the approximation: changing it changes how the approximator responds to states. Because the function must be differentiable with respect to θ for every state, gradient-based methods can use changes in the parameters to improve predictions.
One Example at a Time
The learning process is described at discrete time steps: t = 0, 1, 2, 3, and so on. At time t, the learner observes an example consisting of a state Sₜ and its true value under the policy, written vπ(Sₜ). The state example may be randomly selected or may come from successive interaction with the environment.
Stochastic gradient descent uses a sequence of parameter updates based on examples encountered over time. Rather than treating the approximator as a table of exact values, it repeatedly adjusts the shared parameter vector as examples arrive. This makes the approach particularly suited to online reinforcement learning, where learning can proceed from successive experience.
What do you think happens?
At the next time step, does the learner continue using exactly the same parameter vector?
Reveal answer
Answer: No, the next step uses an updated vector.
The prediction at one step uses θₜ. After the parameter update, the next step uses θₜ₊₁.
Reading the Time-Indexed Weights
The symbol θₜ means the current parameter vector at time step t. The symbol θₜ₊₁ means the parameter vector used at the next time step, after the current example has produced an update. The subscript identifies when that version of the vector is being used; it does not identify a different kind of parameter.
Following One Parameter Update
An example at time t contains state Sₜ and its true policy value vπ(Sₜ). Which parameter vector produces the current prediction, and which vector is used next?
Current prediction: The approximator evaluates Sₜ using the current vector θₜ, producing v̂(Sₜ, θₜ).
Parameter adjustment: The learning method adjusts the parameter vector using the observed example.
Next time step: The updated vector is denoted θₜ₊₁ and is used for the next prediction.
The notation records a sequence: θₜ is used before the update, and θₜ₊₁ is used after the update.
Why Exact Fits Are Not the Goal
A correct target value for an observed state does not guarantee that the approximator will reproduce it exactly. The approximator has limited resources and limited resolution. In general, there may be no single θ that makes every state correct, and there may not even be a θ that makes every observed example exactly correct.
Imagine that an observed state has a correct policy value of 8, while the current approximator predicts 5. The example indicates that the prediction should be improved. An update changes the shared parameter vector in response to this example, but the available parameterization may prevent the new prediction from becoming exactly 8. The learning objective is useful approximation across states, not a guaranteed perfect fit for every example.
Generalization Beyond Observed States
The parameters are shared by the approximate value function rather than belonging to one permanently isolated state entry. Consequently, adjusting the weight vector from an observed example changes the behavior of the function that is evaluated at other states as well. This is the basis of generalization: the approximator should provide values for states that have not yet appeared in the examples.
Suppose the learner observes one state and uses its policy value to adjust θ. The updated vector can then be supplied to v̂ for another state that has not yet been observed. That second state receives a prediction from the same approximate value function, even though no separate exact answer for it was stored.
Common Interpretation Mistakes
Treating the weight vector as a table of one exact value for every state.
The weight vector is a finite collection of real-valued parameters that controls an approximate value function.
Fix:
Interpret θ as the adjustable parameter set used by v̂(s, θ) to produce estimates.Assuming θₜ and θₜ₊₁ are unrelated parameter sets.
The notation records successive versions of the parameter vector during learning.
Fix:
Read θₜ as the vector at time t and θₜ₊₁ as the updated vector used at the next step.Expecting every observed target to be matched exactly.
Limited resources and limited resolution may make a perfect fit for every state or example impossible.
Fix:
Evaluate the method as an approximation intended to produce useful values across states.Assuming an unobserved state cannot receive a value estimate.
Generalization is part of the purpose of function approximation.
Fix:
Use the learned parameter vector with the approximate value function to produce predictions for other states.
Practice the Update Story
Explain the following sequence in your own words: at time t, the learner has θₜ, observes Sₜ with its true policy value, produces an approximate prediction, and then proceeds to the next time step. Your explanation should identify what changes, what remains the same, and why the next prediction can also be made for a state not yet observed.
Hints
- The state example and its true value provide information for the parameter update.
- The next parameter vector is called θₜ₊₁.
- The same approximate value function can be evaluated at other states.
When reading an explanation of stochastic or semi-gradient value prediction, track three items separately: the observed example, the current parameter vector, and the prediction produced by the approximator. This prevents the common confusion between a target value, an estimate, and the parameters that control the estimate.
Key Takeaways
- Stochastic gradient descent adjusts a parameterized approximate value function through a sequence of observed examples.
- The weight vector is a finite set of real-valued parameters controlling the approximate value function v̂(s, θ).
- θₜ denotes the parameters at time t, while θₜ₊₁ denotes the updated parameters used at the next step.
- A correct target does not guarantee an exact prediction because the approximator has limited resources and resolution.
- Generalization allows the learned approximator to provide predictions for states that have not yet appeared in the examples.
Key Takeaways
- Value prediction can use a differentiable function approximator instead of storing a separate exact value for every state.
- The weight vector contains the adjustable parameters that determine the approximator's predictions.
- Learning moves from θₜ to θₜ₊₁ as observed state examples produce parameter updates.
- Approximation seeks useful predictions across states, not necessarily exact agreement with every observed target.
- Because parameters are shared by the approximator, observations can support predictions for states not yet observed.