Semi-Gradient Q-Learning
Semi-Gradient Expected Sarsa is a one-step algorithm for action values.
The Three-Question Trace
When studying Semi-Gradient Q-Learning in this context, begin with the algorithm described as Semi-Gradient Expected Sarsa. A reliable mental trace has three questions: What kind of quantity is being learned? Which target policy is involved? Is importance sampling used? The answers identify the method and explain when it matches Semi-Gradient Q-Learning.
Action Values in One Step
Semi-Gradient Expected Sarsa is a one-step algorithm for action values. The phrase action values identifies what the method is concerned with, while one-step identifies the algorithmic scope emphasized by the source.
The method is best understood by keeping its classification separate from its policy choice. First, it is a one-step action-value algorithm. Next, a target policy π is considered. The important comparison is whether that target policy is greedy with respect to the current action-value function. Finally, the method has a specific sampling choice: it does not use importance sampling.
When Greed Makes the Methods Match
A Greedy Target Policy
Determine the relationship between Semi-Gradient Expected Sarsa and Semi-Gradient Q-Learning when the target policy π is greedy with respect to the current action-value function.
Identify the target policy: The target policy π is specified as greedy with respect to the current action-value function.
Compare the methods: Under this policy condition, the source states that Semi-Gradient Expected Sarsa is equivalent to Semi-Gradient Q-Learning.
Record the sampling choice: The equivalence does not change the stated sampling fact: Semi-Gradient Expected Sarsa does not use importance sampling.
With a target policy greedy with respect to the current action-value function, Semi-Gradient Expected Sarsa and Semi-Gradient Q-Learning are equivalent.
The important condition is not merely that a policy exists. It is that π is greedy with respect to the current action-value function. Once that condition holds, the source identifies the two methods as equivalent.
The Importance-Sampling Decision
Semi-Gradient Expected Sarsa does not use importance sampling. This is a defining sampling choice for the method described in the source pack.
The absence of importance sampling gives the algorithm a clear description: it does not add an importance-sampling correction. The source also warns against turning that fact into a universal rule for every setting. Whether to omit importance sampling is straightforward in the tabular case, but it remains a judgment call when function approximation is involved.
A Compact Parameter Trace
The source distinguishes the tabular case from function approximation. In the tabular case, omitting importance sampling is described as clearly appropriate. With function approximation, the source describes the same omission as a judgment call. The practical trace is therefore not a calculation but a classification: first identify the representation setting, then record the appropriate level of confidence about the sampling choice.
Use precise wording in explanations: say that omission of importance sampling is clearly appropriate in the tabular case, while the choice remains a judgment call with function approximation. This preserves the distinction made by the source.
Check Your Understanding
A learner makes four claims: Semi-Gradient Expected Sarsa is a multi-step method; it learns action values; it becomes equivalent to Semi-Gradient Q-Learning under a greedy target policy; and it always requires importance sampling. Identify the two correct claims and rewrite the two incorrect claims.
Hints
- Recall the number of steps associated with Semi-Gradient Expected Sarsa in the source.
- Check the condition placed on the target policy π.
- Separate the method's actual sampling choice from broader decisions that may depend on the setting.
What do you think happens?
What should you conclude if the target policy π is greedy with respect to the current action-value function?
Reveal answer
Answer: The methods are equivalent.
The source states that Semi-Gradient Expected Sarsa is equivalent to Semi-Gradient Q-Learning when the target policy is greedy with respect to the current action-value function.
Treating Semi-Gradient Expected Sarsa and Semi-Gradient Q-Learning as unrelated methods in every situation.
The source gives a specific condition under which they are equivalent: the target policy π is greedy with respect to the current action-value function.
Fix:
Check the target policy before deciding whether the methods differ.Saying that Semi-Gradient Expected Sarsa uses importance sampling.
The source explicitly states that the algorithm does not use importance sampling.
Fix:
Describe the method as omitting importance sampling.Turning the tabular-case conclusion into a universal rule for function approximation.
The source calls the omission clearly appropriate in the tabular case but a judgment call when function approximation is involved.
Fix:
State the representation setting and preserve the distinction between clear appropriateness and judgment.Forgetting what the algorithm learns.
The source identifies it as a one-step algorithm for action values.
Fix:
Begin the explanation by identifying it as a one-step action-value algorithm.
Key Takeaways
- Semi-Gradient Expected Sarsa is a one-step algorithm for action values.
- It is equivalent to Semi-Gradient Q-Learning when the target policy π is greedy with respect to the current action-value function.
- Semi-Gradient Expected Sarsa does not use importance sampling.
- Omitting importance sampling is clearly appropriate in the tabular case.
- With function approximation, whether to omit importance sampling remains a judgment call.
Key Takeaways
- Semi-Gradient Expected Sarsa is a one-step action-value algorithm.
- A greedy target policy makes Semi-Gradient Expected Sarsa equivalent to Semi-Gradient Q-Learning.
- The method does not use importance sampling.
- The appropriateness of omitting importance sampling depends on the setting: it is clear in the tabular case and a judgment call with function approximation.