Concepts / Semi-gradient Sarsa

Semi-gradient Sarsa

Calculate the n-Step TD Error before applying the semi-gradient Sarsa update.

  • Programming

The Two-Stage Learning Step

Semi-gradient Sarsa separates each learning step into two connected stages. First, it calculates the n-Step TD Error. Next, it applies the usual semi-gradient Sarsa update to the value-function weights θ. The error calculation supplies the learning signal for the update, while the update changes the weights.

preparesupply errorExperiencereward and next staten-Step TD Errorlearning signalWeight updatechange θ
What happens first, and why must the n-Step TD Error be calculated before the value-function weights are updated?

Following One Transition

Imagine the algorithm has a current state and action. It takes that action, observes a reward R and a next state S′, and then determines whether S′ is terminal. The terminal decision selects the update path. Before the selected semi-gradient Sarsa update is applied, the algorithm calculates the n-Step TD Error for the step.

take Atest S′terminalnonterminalCurrentstate-actionS and AReward and next stateR and S′Terminal testTerminal updatereward target; end episodeNonterminal updatechoose A′; includeestimated value
When a transition reaches a terminal state, how does control flow differ from the nonterminal update path?

A Structural Trace

Trace one transition without using numerical values.

Observe: The current action is taken, and the algorithm receives reward R and next state S′.

Classify: The algorithm tests whether S′ is terminal.

Prepare: The n-Step TD Error is calculated before the semi-gradient Sarsa update.

Update: The selected terminal or nonterminal rule is used to update the value-function weights θ.

Continue or stop: A terminal transition ends the episode. A nonterminal transition continues with S′ and a selected next action A′.

The update is not the first operation after observing the transition. Error calculation comes first, and the terminal test determines which update path follows.

Why the Error Comes First

The n-Step TD Error represents the point at which the algorithm has prepared its learning signal for the next part of the procedure. The semi-gradient Sarsa update uses that result to change the learning system. Treating the stages as separate makes the dependency clear: the update follows the error calculation because the calculated error is the input that prepares the update.

This source describes the relationship structurally rather than numerically. It does not provide the numerical form of the n-Step TD Error equation or a worked numerical update. Therefore, the important lesson here is the control-flow dependency, not a particular arithmetic calculation.

When tracing an implementation or algorithm description, mark the end of error calculation before marking the beginning of the parameter update. This prevents the two connected stages from being mistaken for one indistinguishable operation.

comparecompareguideCurrent estimateMulti-step returnn-Step TD Errorlearning signalSemi-gradient updatechange θ
How does the n-Step TD Error represent the difference between the current estimate and the multi-step return used to guide learning?

Episode-Level Weight Changes

Episodic Semi-gradient Sarsa learns a value function one episode at a time. It begins with arbitrarily initialized value-function weights θ. During an episode, it obtains an initial state and action, repeatedly takes the current action, observes a reward and next state, and updates θ. Each transition therefore gives the algorithm another opportunity to adjust its value-function representation.

For a nonterminal transition, the algorithm chooses a next action from q̂(S′, ·, θ), includes the estimated value associated with that next state and action, and then advances the current state and action to S′ and A′. For a terminal transition, the observed reward supplies the target component and the episode ends instead of advancing to a next action.

take current actionreturn R and S′calculateguide updateprovide updated estimatesEpisodeEnvironmentn-Step TD ErrorWeights θupdated value function
How do rewards, visited state-action features, the n-Step TD Error, and successive parameter updates change the weights from one step to the next?
Transition typeInformation usedWhat happens next
TerminalObserved reward RUse the terminal update rule and end the episode
NonterminalReward R, next state S′, and estimated value after choosing A′Use the nonterminal update rule and continue with S′ and A′

The terminal test selects the update path.

Terminal Transitions

When the next state S′ is terminal, the algorithm follows the terminal update path. The observed reward is used as the target component, the value-function weights θ are updated through the semi-gradient Sarsa procedure after the n-Step TD Error is calculated, and the episode ends.

EASY

A transition produces reward R and a next state S′. The terminal test says that S′ is terminal. Describe the control flow in the correct order.

Hints
  • Start with the observed reward and terminal decision.
  • Place n-Step TD Error calculation before the weight update.
  • Do not add a next action or continue the episode.

Nonterminal Transitions

When S′ is nonterminal, the algorithm chooses a next action A′ from q̂(S′, ·, θ). The nonterminal target path includes the estimated value associated with that next state and action. After the n-Step TD Error has been calculated and the semi-gradient update has been applied, the current state and action advance to S′ and A′ so the episode can continue.

This is the main difference from the terminal path: a nonterminal transition prepares another state-action pair for the next step. The terminal path stops after its update, while the nonterminal path updates θ and then continues with the newly selected action.

When checking a nonterminal trace, verify three items in order: R and S′ were received, A′ was chosen from q̂(S′, ·, θ), and S and A were advanced to S′ and A′ after the update.

Tracing Errors in Control Flow

  • Applying the semi-gradient Sarsa update before calculating the n-Step TD Error.

    The source treats the error calculation as the stage that prepares the following update.

    Fix: Calculate the n-Step TD Error first, then apply the semi-gradient Sarsa update.

  • Treating terminal and nonterminal transitions as the same path.

    Terminal transitions use the observed reward as the target component and end the episode.

    Fix: Use the terminal update path and stop. Choose A′ only on the nonterminal path.

  • Choosing A′ without checking whether S′ is terminal.

    The next action is part of the nonterminal path, not the terminal one.

    Fix: Perform the terminal test first. Choose A′ only when S′ is nonterminal.

  • Advancing the current state and action at the wrong time.

    The nonterminal path advances to S′ and A′ after the update, while the terminal path ends the episode.

    Fix: Keep the current state-action pair through the update, then advance only on the nonterminal path.

If a traced result differs from the expected algorithm, find the first divergence in the control flow. Check the values immediately after the action, R and S′. Then check the terminal decision, the target path selected by that decision, and what happens after the update. For a nonterminal transition, also check that A′ was chosen from q̂(S′, ·, θ) before the nonterminal update and that S and A were then advanced to S′ and A′.

Practice the Update Order

MEDIUM

Write a five-step trace for an episodic semi-gradient Sarsa transition. Include the current state and action, the observed reward and next state, the terminal test, the n-Step TD Error calculation, and the selected weight-update path. Then write the different final step for a nonterminal transition.

Hints
  • The error calculation must appear before the semi-gradient update.
  • A terminal transition uses the observed reward as the target component and ends the episode.
  • A nonterminal transition chooses A′ from q̂(S′, ·, θ), includes the estimated value, and continues with S′ and A′.

What do you think happens?

A transition reaches a terminal state. Should the algorithm choose a next action before applying the update?

  • Yes, every transition requires a next action
  • No, the terminal path uses the observed reward and ends the episode
  • Only after the weights are updated
Reveal answer

Answer: No, the terminal path uses the observed reward and ends the episode.

The terminal test selects the terminal update rule. Choosing A′ and continuing with S′ are parts of the nonterminal path.

Key Takeaways

  1. Semi-gradient Sarsa has two connected stages: calculate the n-Step TD Error, then update the value-function weights θ.
  2. The n-Step TD Error prepares the learning signal used by the following semi-gradient Sarsa update.
  3. Episodic learning processes one episode at a time, repeatedly changing θ after transitions.
  4. A terminal transition uses the observed reward as the target component and ends the episode.
  5. A nonterminal transition chooses A′ from q̂(S′, ·, θ), includes the estimated value, and continues with S′ and A′.

Key Takeaways

  • Calculate the n-Step TD Error before applying the semi-gradient Sarsa update.
  • Use the terminal test to select between the terminal and nonterminal update paths.
  • Terminal transitions use the observed reward as the target component and end the episode.
  • Nonterminal transitions choose A′, include its estimated value, and continue with S′ and A′.
  • Episodic Semi-gradient Sarsa changes the value-function weights θ repeatedly as it processes transitions.