Concepts / Differential Semi-gradient Sarsa

Differential Semi-gradient Sarsa

The differential n-step return is G(n)t = R̄ + V̄(S(t+n)) − V̄(S(t)).

  • Programming

Why the Return Is Differential

Differential Semi-gradient Sarsa learns from continuing transitions. Instead of treating learning as a finite episode with a final endpoint, it updates an estimate of average reward together with value-function weights. For function approximation, the useful n-step target is therefore written in differential form: it combines an estimate of average reward with a difference between two estimated state values.

The differential n-step return is the starting point for constructing the n-step TD error. Form the return first; define the error from that return afterward.

thenthenafter errorafter errorafter updatesafter updatesReward and nextstateobserve transitionNext actionselect actionTD errorcalculate before updatesAverage rewardestimateupdate R̄Value-functionweightsupdate θState-actionassignmentassign S and A
What happens first, next, and last when a transition is processed?

Reading the Differential Return

G(n)t = R̄ + V̄(S(t+n)) − V̄(S(t))

Read the expression from left to right. R̄ is an estimate of η(π), the average reward associated with the policy. V̄(S(t+n)) is the estimated value of the later state reached n steps after time t. V̄(S(t)) is the estimated value of the starting state at time t. The later-state estimate is added, and the starting-state estimate is subtracted. The result is the n-step return in differential form.

R̄average reward estimate+addV̄(S(t+n))later-state value−subtractV̄(S(t))starting-state value
What does each term contribute to G(n)t, and how are the two value estimates compared?

From Return to Error

A Symbolic Three-step Trace

Suppose the chosen step length is n = 3. Identify the later state used in the differential return and describe the order in which the return and TD error are formed.

Select the later state: With n = 3, the later-state term is V̄(S(t+3)). The structure changes with n because the later state is three steps after time t.

Form the return: Use G(3)t = R̄ + V̄(S(t+3)) − V̄(S(t)). The average reward estimate and the two state-value estimates are combined before the TD error is defined.

Form the TD error: The n-step TD error is defined using this differential n-step return. The return and the error are related quantities, but they are not the same named quantity.

Check the range: The stated expression is used when n is at least 1 and t+n is less than T. For this trace, that means t+3 must be less than T.

For n = 3, use the value estimate at S(t+3), form the differential return, and then define the n-step TD error from that return.

The important distinction is procedural. The differential return supplies the target quantity used to define the n-step TD error. The error must be calculated before either learned quantity is updated. This separation makes it possible to inspect whether the return was formed correctly before checking the subsequent updates.

used to defineG(n)tdifferential returnδ(n)tn-step TD error
How is the n-step TD error obtained from the differential return?

Two Quantities Updated

Differential Semi-gradient Sarsa updates two learned components. The value-function weights control the differentiable approximation q_hat, which estimates state-action values. The average reward estimate R̄ represents the algorithm's estimate of η(π). Together, these components support the next prediction and the next update in a continuing-transition setting.

controlscontributes tovalue estimates contributedefinesupdatesupdatesθvalue-function weightsR̄average reward estimateq_hatstate-action estimateG(n)tdifferential returnTD errordrives the step
How do the value-function weights and average reward estimate contribute to later predictions and updates?

The two updates are not interchangeable with the TD-error calculation. First calculate the error; only afterward update the average reward estimate and the value-function weights.

Tracing an Algorithm Step

  1. Take an action and observe the reward and next state.
  2. Select the next action.
  3. Calculate the temporal-difference error, using the differential return that supplies its target.
  4. Update the average reward estimate R̄.
  5. Update the value-function weights θ.
  6. Assign the next state and action for the following step.

This order is an implementation checkpoint. The error must be available before either learned quantity changes. The state and action assignment also comes last in the trace because replacing them too early can cause the weight update to use the wrong state-action pair.

When tracing a step by hand, record delta first, then R_bar, then theta, and finally the next state-action assignment. Comparing your trace with that sequence makes the first divergence easier to locate.

Finding the First Divergence

  • Updating the average reward estimate before calculating the TD error.

    The stated procedure calculates the temporal-difference error before either learned quantity is updated.

    Fix: Calculate the error first, then update R_bar and the value-function weights.

  • Treating the differential return and the n-step TD error as the same quantity.

    The differential return supplies the target used to define the n-step TD error; they are related but differently named quantities.

    Fix: Write down G(n)t first, then form the n-step TD error from it.

  • Replacing the state and action before the weight update.

    The weight update may then use the wrong state-action pair.

    Fix: Complete the error calculation and learned-quantity updates before making the next state-action assignment.

  • Using the wrong later state in the return.

    Changing n changes which later state appears in the expression.

    Fix: Check that the later estimate is V̄(S(t+n)) and that t+n is less than T.

Practice Trace

MEDIUM

A trace uses n = 2. It forms the differential return with V̄(S(t+2)), updates R_bar, calculates the TD error, updates theta, and then assigns the next state and action. Identify the first step that is out of order and explain what should happen instead.

Hints
  • The TD error must be calculated before either learned quantity is updated.
  • The return uses the later state S(t+n), so for n = 2 the later state is S(t+2).
  • The state-action assignment belongs after the learned-quantity updates.

What do you think happens?

In the trace, which checkpoint should be inspected first when the final result differs from the expected result?

  • The final value-function weights
  • The first calculation of the TD error
  • The last state-action assignment
Reveal answer

Answer: The first calculation of the TD error

The debugging procedure checks the calculation in order: first the terms in the error, then the average reward estimate update, then the value-function weight update, and finally the state-action assignments. The earliest mismatch identifies the first divergence.

Key Takeaways

  1. For function approximation in a continuing setting, the n-step return is written as G(n)t = R̄ + V̄(S(t+n)) − V̄(S(t)).
  2. R̄ estimates η(π), while the two V̄ terms estimate the later and starting states.
  3. The displayed differential return applies when n is at least 1 and t+n is less than T.
  4. The n-step TD error is defined from the differential return and must be calculated before learned quantities are updated.
  5. A reliable trace checks the error, R_bar, theta, and the next state-action assignment in that order.

Key Takeaways

  • Differential Semi-gradient Sarsa learns from continuing transitions by updating both an average reward estimate and value-function weights.
  • The differential n-step return compares the estimated value of a later state with the estimated value of the starting state, while also including the average reward estimate.
  • The return must be formed before the n-step TD error, and the error must be calculated before either learned quantity is updated.
  • The safest debugging method is to compare the calculation checkpoints in order and stop at the first mismatch.