Reinforcement learning episodes
Episodic Semi-gradient Sarsa learns a value function by updating value-function weights θ during episodes.
The Episode as a Learning Loop
Episodic Semi-gradient Sarsa learns a value function one episode at a time. It begins with arbitrarily initialized value-function weights θ. During an episode, the algorithm starts with an initial state and action, takes that action, observes a reward and a next state, and then updates θ. The next decision is determined by whether the next state is terminal.
The Two Update Paths
The terminal test is the control-flow decision that selects the semi-gradient update rule. After the current action produces a reward R and next state S′, the algorithm checks S′. If S′ is terminal, the terminal path uses the observed reward as the target component and ends the episode. If S′ is nonterminal, the algorithm chooses a next action A′ using the estimated values q̂(S′, ·, θ), includes the estimated value of that next state-action choice in the target, updates θ, and continues with S′ and A′.
How Weights Move Through a Trace
A Three-Transition Episode
Trace the control flow and weight changes for an episode with two nonterminal transitions followed by one terminal transition.
Episode begins: Start with the arbitrarily initialized weights θ and obtain an initial state and action.
First transition: Take the current action and observe R and S′. If S′ is nonterminal, choose A′ from q̂(S′, ·, θ), include the estimated next state-action value in the target, and update θ. The current state and action then advance to S′ and A′.
Second transition: Repeat the same nonterminal path if the newly observed next state is still nonterminal. This produces another update to θ before the episode reaches its end.
Final transition: When the observed next state is terminal, use the terminal path. The observed reward supplies the target component, θ is updated, and the episode ends instead of advancing to a next action.
The weight sequence is initial θ, θ after the first update, θ after the second update, and θ after the terminal update. The terminal transition changes the update path and ends the episode.
The trace shows what can be concluded without assigning invented numeric values to θ. Every processed transition can update θ. A nonterminal transition also produces the next state-action pair used to continue the episode. A terminal transition produces an update and then stops the episode, so there is no subsequent S′ and A′ advancement within that episode.
From Estimates to Weights
The update uses information available immediately after the action: the observed reward R and the next state S′. On the nonterminal path, the algorithm first chooses A′ using q̂(S′, ·, θ), then includes the estimated value for that next state-action choice in the target before changing θ. On the terminal path, the observed reward is the target component and no next action is selected. Thus, the terminal test determines which information contributes to the update and whether the episode continues.
The Terminal Boundary
A terminal transition is the boundary of an episode. When S′ is terminal, the algorithm uses the observed reward as the target component, updates θ, and ends the episode. It does not choose A′ or advance the working state-action pair to S′ and A′. The source description specifies the terminal target component and the end of the episode; it does not assign a separate numeric value to a post-terminal state.
Tracing Errors in Control Flow
Choosing A′ before checking whether S′ is terminal.
The terminal test determines whether the nonterminal path is used. Terminal transitions end the episode and do not choose a next action.
Fix:
Inspect S′ first. Choose A′ only on the nonterminal path.Using the nonterminal target path for a terminal transition.
Terminal transitions use the observed reward as the target component and end the episode.
Fix:
Select the terminal update rule when S′ is terminal.Advancing to S′ and A′ after a terminal update.
The terminal transition ends the episode.
Fix:
Stop the episode immediately after the terminal update.Inspecting only the final weights when a trace is wrong.
A later mismatch may be caused by the first incorrect terminal decision, target choice, or state-action advancement.
Fix:
Find the first divergence. Check R and S′, the terminal decision, the selected target expression, and the destination after the update.
Control-Flow Practice
A transition produces reward R and next state S′. You determine that S′ is nonterminal. What must happen before θ is updated, and what state-action information is used to continue the episode?
Hints
- The nonterminal path requires a next action.
- The next action is chosen using q̂(S′, ·, θ).
- After the update, the continuing state-action pair is S′ and A′.
A transition produces reward R and a terminal next state S′. Which update path should be selected, and what happens immediately after θ is updated?
Hints
- The terminal test selects the update rule.
- The observed reward is the target component on this path.
- The episode ends rather than advancing to a next action.
Episode-Level Takeaways
- Episodic Semi-gradient Sarsa learns a value function by updating θ while processing one episode at a time.
- After each action, the algorithm observes R and S′ and then performs the terminal test.
- A nonterminal transition chooses A′ using q̂(S′, ·, θ), includes the next estimate, updates θ, and continues with S′ and A′.
- A terminal transition uses the observed reward as the target component, updates θ, and ends the episode without choosing A′.
- When debugging a trace, locate the first control-flow divergence before comparing final weights.
Key Takeaways
- The terminal test is the central control-flow decision in an episodic semi-gradient Sarsa update.
- Nonterminal transitions use a next action and its estimated value to continue learning through the episode.
- Terminal transitions use the observed reward as the target component and end the episode.
- Tracing θ means recording the update after each transition and checking whether the episode continues or stops.
- The first place to debug is the earliest mismatch in the observed values, terminal decision, target path, or state-action advancement.