Concepts / Semi-Gradient TD(0)

Semi-Gradient TD(0)

SGD variations are used to find a good weight vector for function approximation.

  • Programming

From Samples to Weights

When a value function is approximated, learning involves finding a good weight vector. Variations of stochastic gradient descent are used for this purpose: rather than trying to determine the best weight vector in one step, the method uses sampled experience to produce successive updates.

Semi-gradient TD(0) belongs to a broader family called n-step semi-gradient TD. The family is useful in the on-policy case with a fixed policy because it provides a way to update the approximate value function from sampled experience while using a target that may include both observed information and an estimate. The choice of n determines how far the method looks through the sampled experience before forming its target.

One Update in Motion

A semi-gradient TD(0) update can be viewed as a short information flow. The current state estimate, the next-state estimate, and the reward are combined into a one-step target. The difference between the current estimate and that target is the temporal-difference error. The weight vector is then updated in response to that error.

compared withcontributes tocontributes toforms comparisonforms comparisondrives updateCurrent stateestimateapproximate valueOne-step targetcurrent reward and nextestimateTD errorestimate versus targetWeight vectorupdated parametersNext-state estimateapproximate valueRewardsampled information
How do the current estimate, next-state estimate, reward, TD error, and weight vector participate in one update?

Tracing a One-Step Update

Trace the conceptual role of each quantity in one semi-gradient TD(0) update.

Start with the current estimate: The approximate value associated with the current state is available from the current weight vector.

Form a one-step target: The immediate reward and the estimate of the next state are used together to form the target for this update.

Measure the TD error: The target is compared with the current estimate. Their difference is the temporal-difference error.

Update the weights: The weight vector is adjusted in response to the TD error, producing a new approximation for later updates.

The update moves the approximate value function using one sampled transition rather than waiting for a complete episode.

The N-Step Family

N-step semi-gradient TD provides a unifying view of several learning methods. At one step, the target is the short, bootstrapped target used by semi-gradient TD(0). As n increases, the target incorporates more of the sampled sequence before relying on an estimate. At the infinity-step end, the method becomes gradient Monte Carlo. Thus, semi-gradient TD(0) and gradient Monte Carlo are special cases of the same n-step family.

increase nincrease nn = 1semi-gradient TD(0)Intermediate nmulti-step returnn = infinitygradient Monte Carlo
How does changing n move the method from a one-step bootstrapped estimate toward a full episode return?
Position in the familyMethodTarget character
One-step caseSemi-gradient TD(0)One-step bootstrapped estimate
Intermediate stepsIntermediate n-step semi-gradient TDMulti-step return
Infinity-step caseGradient Monte CarloFull episode return

The n-step family connects semi-gradient TD(0) and gradient Monte Carlo.

This family view explains why n-step semi-gradient TD is natural for an on-policy fixed-policy problem. The same basic learning idea can use a short target, a longer multi-step target, or the full episode return, depending on the chosen value of n.

Where the Gradient Stops

The name semi-gradient is important. These methods use a weight vector in the update target, so the target depends on the current approximation. However, when the update direction is computed, that dependence of the target on the weight vector is ignored. Because part of the dependency is omitted, the method is not a true gradient method.

would require target dependencedependence omittedTD targetdepends on weightsTD targetweight dependence retainedUpdate directiongradient calculationUpdate directiontarget dependence ignored
Which dependency is present in the TD target but omitted when the semi-gradient update direction is computed?

Common Misreadings

  • Treating semi-gradient TD as a true gradient method.

    The target's dependence on the weight vector is ignored when the update direction is computed.

    Fix: Call it a semi-gradient method: the target uses the weight vector, but that target dependence is not included in the gradient.

  • Treating semi-gradient TD(0) and gradient Monte Carlo as unrelated algorithms.

    Both are special cases of n-step semi-gradient TD.

    Fix: Use the n-step family as the organizing view: semi-gradient TD(0) is the one-step case and gradient Monte Carlo is the infinity-step case.

  • Assuming that increasing n leaves the target unchanged.

    Changing n changes the target from a one-step bootstrapped estimate toward a multi-step return and, at the infinity-step end, a full episode return.

    Fix: Track n explicitly when comparing methods in the family.

Check Your Understanding

MEDIUM

Explain in your own words why semi-gradient TD(0) is called semi-gradient rather than true gradient. Then place semi-gradient TD(0) and gradient Monte Carlo at the correct ends of the n-step family.

Hints
  • Focus on the part of the TD target that depends on the current weight vector.
  • Use one-step for semi-gradient TD(0) and infinity-step for gradient Monte Carlo.

Answer Check

Give the two essential distinctions required by the practice prompt.

Explain the name: Semi-gradient TD uses a weight vector in its target, but ignores the target's dependence on that weight vector when computing the gradient.

Locate the endpoints: Semi-gradient TD(0) is the one-step special case of n-step semi-gradient TD, while gradient Monte Carlo is the infinity-step special case.

The method is semi-gradient because the gradient calculation omits part of the target's weight dependence, and the n-step family connects the one-step and infinity-step cases.

Key Takeaways

  1. Variations of stochastic gradient descent are used to find a good weight vector for function approximation.
  2. N-step semi-gradient TD is suited to the on-policy setting with a fixed policy.
  3. Semi-gradient TD(0) is the one-step special case of n-step semi-gradient TD.
  4. Gradient Monte Carlo is the infinity-step special case.
  5. Semi-gradient TD methods are not true gradient methods because they ignore the target's dependence on the weight vector when computing the gradient.

Key Takeaways

  • Stochastic-gradient-descent variations update a weight vector incrementally to improve a function approximation.
  • N-step semi-gradient TD unifies one-step, multi-step, and infinity-step targets for the on-policy fixed-policy setting.
  • Semi-gradient TD(0) is the one-step case, while gradient Monte Carlo is the infinity-step case.
  • Semi-gradient methods are not true gradient methods because they omit the target's dependence on the current weight vector from the gradient calculation.