Concepts / Episodic Semi-gradient Sarsa for Control

Episodic Semi-gradient Sarsa for Control

A learning curve is a way to represent performance over time.

  • Programming

From Curves to Control

Episodic Semi-gradient Sarsa can be understood from two connected perspectives. Inside an episode, the algorithm repeatedly updates value-function weights θ after observing what happens when it takes an action. Across many episodes, its performance can be displayed as a learning curve. The first perspective explains how learning happens; the second helps us study how performance changes over time.

The central control-flow question is whether the next state is terminal. That decision determines which update rule is used.

One Episode of Learning

Episodic Semi-gradient Sarsa processes one episode at a time. It begins with arbitrarily initialized value-function weights θ. For an episode, it obtains an initial state and action. It then takes the current action, observes a reward and a next state, and updates θ. The algorithm repeats this process until the episode ends.

use current estimatestake Ainspect S′select ruleupdate θif nonterminalθinitial weightsS, Atake current actionR, S′observed resultterminal testchoose update pathtargetreward or reward with nextvalueθupdated weightsS′, A′next nonterminalstate-action pair
How does each state-action transition provide information for changing θ before the next transition?

The diagram shows the repeated structure of the algorithm. The current state-action pair produces an observed reward and next state. The terminal test then selects the target used to update θ. If the transition is nonterminal, the next state and next action become the current state and action for the next step.

Two Update Paths

The terminal test is not a minor bookkeeping detail. It determines whether the update uses only the observed reward or also includes an estimated value for the next state-action pair.

observe S′yesnouse rewardfinish episodeinclude next valueadvanceS, Acurrent pairS′terminal?terminalreward targetnonterminalchoose A′θ updateobserved rewardθ updatereward and estimated nextvalueS′, A′continue episodeepisode endstop
What update path does the algorithm follow when S′ is nonterminal, and what changes when S′ is terminal?
Transition resultTarget informationWhat happens next
TerminalThe observed reward is the target componentUpdate θ and end the episode
NonterminalThe observed reward and the estimated value of the next state-action pair are includedChoose A′, update θ, then continue with S′ and A′

Tracing a Transition

Following One Nonterminal Transition

Suppose the algorithm is processing a current state-action pair S and A. Taking A produces reward R and a next state S′ that is nonterminal.

Observe the result: The algorithm records R and S′ after taking the current action.

Choose the path: Because S′ is nonterminal, the algorithm follows the nonterminal update rule.

Choose the next action: The algorithm chooses A′ using the estimated action values q̂(S′, ·, θ).

Update the weights: The update uses the observed reward together with the estimated value associated with S′ and A′, changing θ.

Advance the stored pair: After the update, S and A are advanced to S′ and A′ so the next transition continues from the new state-action pair.

A nonterminal transition both changes θ and supplies the next state-action pair for continuing the episode.

Following One Terminal Transition

Suppose the current state-action pair S and A produces reward R and a terminal next state.

Observe the result: The algorithm records R and the terminal next state.

Select the terminal rule: The terminal test selects the update whose target component is the observed reward.

Update the weights: The algorithm changes θ using the terminal update.

End the episode: There is no next action to choose and no S′, A′ pair to advance to for another transition.

A terminal transition changes θ and ends processing for that episode.

Mountain Car and Tile Coding

The reported experiment combines semi-gradient Sarsa, tile-coding function approximation, and the Mountain Car example. Tile-coding function approximation is therefore the representation setting in which the value function is learned for this example. The learning process changes the value-function weights θ, and the resulting performance can be examined through learning curves.

use settingrepresent learned valueestimate valuesmeasure performanceMountain Carexample problemtile codingfunction approximationθvalue-function weightsvalue functionestimated action valueslearning curvesperformance over time
How does the Mountain Car setting connect tile-coding function approximation with learned values and measured performance?

The source identifies the components of the experiment but does not provide a detailed tile layout or a numerical example of individual tile activations. The safe conclusion is that tile coding supplies the function-approximation setting for the Mountain Car experiment; the weights θ are updated as semi-gradient Sarsa processes episodes in that setting.

Reading Figure 10.2

A learning curve represents performance over time. In this context, the curves concern semi-gradient Sarsa used with tile-coding function approximation on the Mountain Car example. Figure 10.2 displays several curves so that different step sizes can be considered together.

hold setting, vary parameterhold setting, vary parameterproducesproducessemi-gradient SarsaMountain Car with tilecodingstep sizevalue Astep sizevalue Blearning curveperformance overrepresented periodlearning curveperformance overrepresented period
How should the plotted curves be interpreted without claiming more than the labels and displayed trends support?
  1. Identify the common experimental setting: semi-gradient Sarsa with tile-coding function approximation on the Mountain Car example.
  2. Identify which curve is associated with which step size.
  3. Compare the displayed performance over the period represented in the figure.
  4. Describe only the differences that the plotted curves and their labels support.
  5. Avoid asserting exact numerical values or a universal ranking when the supplied information does not provide them.

Mistakes in Control Flow

  • Treating every transition as nonterminal

    The terminal test determines the update rule, and terminal transitions end the episode.

    Fix: Use the observed reward as the target component, update θ, and stop the episode.

  • Using only the reward on a nonterminal transition

    A nonterminal transition includes the next action's estimated value.

    Fix: Choose A′ using q̂(S′, ·, θ), include its estimated value in the nonterminal target, and then advance to S′ and A′.

  • Choosing A′ before checking whether S′ is terminal

    The terminal decision comes first. A′ belongs to the nonterminal path.

    Fix: Inspect S′, select the terminal or nonterminal rule, and choose A′ only on the nonterminal path.

  • Claiming a precise result from Figure 10.2 without reading its labels

    The source identifies several curves and various step sizes but does not state exact numerical values or a particular ranking.

    Fix: Match each curve to its labeled step size and describe only the displayed comparison.

  • Confusing the function-approximation setting with a detailed tile layout

    The source does not provide that detailed tile-level information.

    Fix: State that tile-coding function approximation is used with semi-gradient Sarsa in the Mountain Car example, without inventing unsupported layout details.

Practice the Decision

MEDIUM

A current state-action pair produces a reward R and a next state S′. First, describe the control-flow checks you perform. Then explain the update path if S′ is terminal. Finally, explain how the path differs if S′ is nonterminal.

Hints
  • The terminal test determines the update rule.
  • The terminal path uses the observed reward as the target component and ends the episode.
  • The nonterminal path chooses A′ using q̂(S′, ·, θ), includes the estimated next value, updates θ, and advances to S′ and A′.
EASY

When interpreting Figure 10.2, write a two-sentence conclusion. In the first sentence, identify what is being varied across the curves. In the second sentence, state what can be concluded from the displayed curves and labels, while avoiding exact numerical or ranking claims not supported by the figure.

Hints
  • The curves are associated with different step sizes.
  • Keep the conclusion tied to the plotted period and the labels.

Key Takeaways

  1. Episodic Semi-gradient Sarsa learns a value function by updating value-function weights θ during episodes.
  2. The terminal test selects between two update paths: a terminal path based on the observed reward and a nonterminal path that also includes the estimated next state-action value.
  3. A nonterminal transition chooses A′, updates θ, and continues with S′ and A′; a terminal transition updates θ and ends the episode.
  4. The Mountain Car experiment combines semi-gradient Sarsa with tile-coding function approximation.
  5. Figure 10.2 compares learning curves associated with different step sizes, so conclusions should remain tied to the labels and trends actually displayed.

Key Takeaways

  • Episodic Semi-gradient Sarsa changes θ as it processes state-action transitions within an episode.
  • The next state's terminal status determines whether the target uses only the observed reward or also an estimated next value.
  • Tile-coding function approximation is part of the Mountain Car setting used to examine semi-gradient Sarsa.
  • Figure 10.2 compares learning curves for different step sizes.
  • A sound interpretation reports only what the plotted curves and their labels support.