Episodic Semi-gradient Sarsa for Control
A learning curve is a way to represent performance over time.
From Curves to Control
Episodic Semi-gradient Sarsa can be understood from two connected perspectives. Inside an episode, the algorithm repeatedly updates value-function weights θ after observing what happens when it takes an action. Across many episodes, its performance can be displayed as a learning curve. The first perspective explains how learning happens; the second helps us study how performance changes over time.
The central control-flow question is whether the next state is terminal. That decision determines which update rule is used.
One Episode of Learning
Episodic Semi-gradient Sarsa processes one episode at a time. It begins with arbitrarily initialized value-function weights θ. For an episode, it obtains an initial state and action. It then takes the current action, observes a reward and a next state, and updates θ. The algorithm repeats this process until the episode ends.
The diagram shows the repeated structure of the algorithm. The current state-action pair produces an observed reward and next state. The terminal test then selects the target used to update θ. If the transition is nonterminal, the next state and next action become the current state and action for the next step.
Two Update Paths
The terminal test is not a minor bookkeeping detail. It determines whether the update uses only the observed reward or also includes an estimated value for the next state-action pair.
| Transition result | Target information | What happens next |
|---|---|---|
| Terminal | The observed reward is the target component | Update θ and end the episode |
| Nonterminal | The observed reward and the estimated value of the next state-action pair are included | Choose A′, update θ, then continue with S′ and A′ |
Tracing a Transition
Following One Nonterminal Transition
Suppose the algorithm is processing a current state-action pair S and A. Taking A produces reward R and a next state S′ that is nonterminal.
Observe the result: The algorithm records R and S′ after taking the current action.
Choose the path: Because S′ is nonterminal, the algorithm follows the nonterminal update rule.
Choose the next action: The algorithm chooses A′ using the estimated action values q̂(S′, ·, θ).
Update the weights: The update uses the observed reward together with the estimated value associated with S′ and A′, changing θ.
Advance the stored pair: After the update, S and A are advanced to S′ and A′ so the next transition continues from the new state-action pair.
A nonterminal transition both changes θ and supplies the next state-action pair for continuing the episode.
Following One Terminal Transition
Suppose the current state-action pair S and A produces reward R and a terminal next state.
Observe the result: The algorithm records R and the terminal next state.
Select the terminal rule: The terminal test selects the update whose target component is the observed reward.
Update the weights: The algorithm changes θ using the terminal update.
End the episode: There is no next action to choose and no S′, A′ pair to advance to for another transition.
A terminal transition changes θ and ends processing for that episode.
Mountain Car and Tile Coding
The reported experiment combines semi-gradient Sarsa, tile-coding function approximation, and the Mountain Car example. Tile-coding function approximation is therefore the representation setting in which the value function is learned for this example. The learning process changes the value-function weights θ, and the resulting performance can be examined through learning curves.
The source identifies the components of the experiment but does not provide a detailed tile layout or a numerical example of individual tile activations. The safe conclusion is that tile coding supplies the function-approximation setting for the Mountain Car experiment; the weights θ are updated as semi-gradient Sarsa processes episodes in that setting.
Reading Figure 10.2
A learning curve represents performance over time. In this context, the curves concern semi-gradient Sarsa used with tile-coding function approximation on the Mountain Car example. Figure 10.2 displays several curves so that different step sizes can be considered together.
- Identify the common experimental setting: semi-gradient Sarsa with tile-coding function approximation on the Mountain Car example.
- Identify which curve is associated with which step size.
- Compare the displayed performance over the period represented in the figure.
- Describe only the differences that the plotted curves and their labels support.
- Avoid asserting exact numerical values or a universal ranking when the supplied information does not provide them.
Mistakes in Control Flow
Treating every transition as nonterminal
The terminal test determines the update rule, and terminal transitions end the episode.
Fix:
Use the observed reward as the target component, update θ, and stop the episode.Using only the reward on a nonterminal transition
A nonterminal transition includes the next action's estimated value.
Fix:
Choose A′ using q̂(S′, ·, θ), include its estimated value in the nonterminal target, and then advance to S′ and A′.Choosing A′ before checking whether S′ is terminal
The terminal decision comes first. A′ belongs to the nonterminal path.
Fix:
Inspect S′, select the terminal or nonterminal rule, and choose A′ only on the nonterminal path.Claiming a precise result from Figure 10.2 without reading its labels
The source identifies several curves and various step sizes but does not state exact numerical values or a particular ranking.
Fix:
Match each curve to its labeled step size and describe only the displayed comparison.Confusing the function-approximation setting with a detailed tile layout
The source does not provide that detailed tile-level information.
Fix:
State that tile-coding function approximation is used with semi-gradient Sarsa in the Mountain Car example, without inventing unsupported layout details.
Practice the Decision
A current state-action pair produces a reward R and a next state S′. First, describe the control-flow checks you perform. Then explain the update path if S′ is terminal. Finally, explain how the path differs if S′ is nonterminal.
Hints
- The terminal test determines the update rule.
- The terminal path uses the observed reward as the target component and ends the episode.
- The nonterminal path chooses A′ using q̂(S′, ·, θ), includes the estimated next value, updates θ, and advances to S′ and A′.
When interpreting Figure 10.2, write a two-sentence conclusion. In the first sentence, identify what is being varied across the curves. In the second sentence, state what can be concluded from the displayed curves and labels, while avoiding exact numerical or ranking claims not supported by the figure.
Hints
- The curves are associated with different step sizes.
- Keep the conclusion tied to the plotted period and the labels.
Key Takeaways
- Episodic Semi-gradient Sarsa learns a value function by updating value-function weights θ during episodes.
- The terminal test selects between two update paths: a terminal path based on the observed reward and a nonterminal path that also includes the estimated next state-action value.
- A nonterminal transition chooses A′, updates θ, and continues with S′ and A′; a terminal transition updates θ and ends the episode.
- The Mountain Car experiment combines semi-gradient Sarsa with tile-coding function approximation.
- Figure 10.2 compares learning curves associated with different step sizes, so conclusions should remain tied to the labels and trends actually displayed.
Key Takeaways
- Episodic Semi-gradient Sarsa changes θ as it processes state-action transitions within an episode.
- The next state's terminal status determines whether the target uses only the observed reward or also an estimated next value.
- Tile-coding function approximation is part of the Mountain Car setting used to examine semi-gradient Sarsa.
- Figure 10.2 compares learning curves for different step sizes.
- A sound interpretation reports only what the plotted curves and their labels support.