Semi-gradient Methods for Value Prediction
SGD performs an immediate, small weight-vector adjustment for each example.
Immediate Weight Adjustments
Semi-gradient methods are easier to understand if you first focus on the timing of an update. A value-prediction system receives examples one after another. Stochastic gradient descent does not wait until every example has been processed before changing its parameters. Instead, it makes a small adjustment to the weight vector immediately after each example. Semi-gradient TD(0) uses this repeated-update idea while estimating the value function of a supplied policy.
The word stochastic refers here to using one example for each update. The update is local because it responds to the example currently being considered, but many such updates can support improvement of an average objective such as mean squared value error, or MSVE.
Gradient Direction and Step Size
Gradient descent minimizes a function by moving in the direction of the negative gradient. The gradient is a vector, not a single number: it has one partial derivative for each component of the weight vector. Each component indicates how the error changes when the corresponding weight changes. Taken together, the gradient points toward the direction of greatest error increase. Subtracting the gradient therefore points toward the direction of most rapid decrease.
The step-size parameter controls the scale of the movement. A positive step size multiplies the descent direction, so it determines how large the adjustment is while the negative gradient determines its direction. The current weight vector is the starting point, and the updated weight vector is the result after taking that scaled step.
| Update approach | When weights change | What supplies the immediate direction |
|---|---|---|
| Ordinary gradient descent | After considering the full objective | The gradient of the objective being minimized |
| Stochastic gradient descent | After each individual example | The example currently being considered |
From Local Error to MSVE
A single stochastic update does not directly examine every example at once. It focuses on the error associated with the one example just considered. This local focus can still improve broader performance because the process repeats across many examples. When those examples follow the same distribution as the states used for the MSVE objective, small corrections on observed examples can have the overall effect of reducing an average measure such as mean squared value error.
Generated example: Imagine a value-prediction system receiving a stream of state examples. After the first example, it makes one small correction. After the next example, it makes another correction based on that new local error. No single correction represents the entire average objective, but the repeated sequence can move performance toward a lower average error when the examples have the relevant distribution.
Keep the two interpretations separate: locally, an update tries to reduce one example's error; broadly, repeated updates across suitable examples work toward reducing an average performance measure. The local step and the overall objective are related, but they are not identical.
Value Estimates from Weights
Semi-gradient TD(0) represents a value function with a differentiable function v̂ that depends on a state and a weight vector θ. The weights are the adjustable parameters. For a current state, the function uses the state together with the current weights to produce an estimated value. After a transition is observed, the algorithm changes θ; using the changed weights can therefore change the estimated value produced for a state.
Purpose and Setup
Semi-gradient TD(0) is used for policy evaluation and value-function estimation. It estimates the value function of a supplied policy π while repeatedly improving the weights of a differentiable value-function approximation.
- The policy π being evaluated
- A differentiable value function v̂ that depends on a state and weight vector θ
- A current weight vector θ
- Episodes generated by the supplied policy
- A current state S from which the next action is selected
One Transition in Order
For one TD(0) transition, preserve the algorithm's order exactly. Begin with the current state S. Choose an action from the policy π(·|S). Then observe the reward R and next state S′. Only after both have been observed can the weight update use the current-state estimate, the next-state estimate, the observed reward, the gradient, and the step size. After the update, assign S to S′ so that the next iteration starts from the observed next state.
Tracing a Single Transition
A trace begins at state S with weight vector θ. The policy selects an action. The environment then supplies reward R and next state S′. Identify the correct role of each item in the update sequence.
Start: Use the current state S and current weight vector θ. The value function supplies the current-state estimate.
Select: Choose an action from the supplied policy π(·|S).
Observe: Receive the reward R and observe the next state S′. These are required before the update can be traced.
Update: Use R, the next-state estimate, the current-state estimate, the gradient, and the positive step-size parameter to adjust θ.
Advance: After the weight update, replace S with S′. The next transition begins from that new current state.
The reliable one-step order is choose an action, observe R and S′, update θ, and then set S to S′.
Episode-Level Tracing
At the episode level, the one-step sequence repeats. The policy chooses an action from the current state, the transition produces R and S′, the algorithm updates θ once for that observed transition, and then the current state is replaced by S′. The episode stops when S′ is terminal. The order matters because the current state used for the update must still be the state from which the action was selected.
The update uses information from both sides of the transition: the current-state estimate describes the state being updated, while the reward and next-state estimate provide the observed transition information used to adjust the weights. The gradient connects that error-related adjustment to the components of θ, and the step size controls the scale.
Tracing Unexpected Results
When a traced result looks unexpected, do not begin by changing the final weight value. Find the first point at which the trace no longer matches the algorithm. Check the action, observations, update inputs, and state assignment in that order.
Using an action that did not come from the supplied policy.
Semi-gradient TD(0) evaluates the supplied policy, so the action-selection step must come from that policy.
Fix:
Verify the action-selection step before checking the numerical update.Updating before observing the reward and next state.
The reward R and next state S′ appear in the update information and must be observed first.
Fix:
Place the observation of R and S′ before the weight update.Using the next state as the current state too early.
The update must use the current S from which the action was selected and the observed S′ as the next state.
Fix:
Update θ first, then set S to S′.Treating the gradient as a single scalar.
The gradient is a derivative vector with one partial derivative for each component of the weight vector.
Fix:
Track the gradient as a vector whose components describe how the error changes with the corresponding weights.Confusing the step size with the update direction.
The negative gradient supplies the descent direction; the positive step size controls the movement's scale.
Fix:
Separate direction, supplied by the negative gradient, from scale, supplied by the step size.
Practice Trace
A trace follows this order: choose an action from π, observe S′ and R, replace S with S′, and then update θ. Identify the first ordering problem and rewrite the one-step order correctly.
Hints
- The reward and next state must be available before the update.
- The update must still use the original current state.
- The state assignment occurs after the weight update.
What do you think happens?
What may happen if a trace assigns S to S′ before applying the update?
Reveal answer
Answer: The update may no longer use the intended current state.
The required order is to observe R and S′, update θ using the current S and observed S′, and only then assign S to S′.
Key Takeaways
- Stochastic gradient descent changes the weight vector immediately after each individual example.
- The gradient points toward greatest error increase, so the negative gradient supplies the direction of decreasing error; the positive step size controls the movement's scale.
- Each update addresses one local example error, while repeated updates across suitable examples can reduce an average objective such as MSVE.
- Semi-gradient TD(0) evaluates a supplied policy by repeatedly improving the weights of a differentiable value-function approximation.
- For one transition, the reliable order is: choose an action, observe R and S′, update θ, and then set S to S′.
Key Takeaways
- SGD makes a small weight adjustment after each individual example rather than waiting for all examples.
- The negative gradient gives the direction of decreasing error, while the step size determines the adjustment's scale.
- Many local example-based updates can support reduction of an average objective such as MSVE.
- Semi-gradient TD(0) is used to evaluate a policy and estimate its value function with differentiable value-function approximation.
- A correct TD(0) trace observes the reward and next state before updating and assigns the next state only after the update.