Action Values in Policy-Gradient Methods
A baseline extends the policy gradient theorem by providing a reference for comparing an action value.
Why Compare with a Reference
A policy-gradient method can use an action value to describe the value associated with the selected action. The theorem can be generalized by comparing that action value with a baseline. The baseline does not replace the action value. It supplies a reference, so the important quantity becomes the action value relative to that reference.
The central comparison is between Q(s,a), which varies with the selected action, and b(s), which must not vary with the action.
What do you think happens?
For one fixed situation s, which quantity should remain the same when the selected action changes?
Reveal answer
Answer: The baseline b(s)
The action value is associated with the selected action, while the baseline is an action-independent reference for the situation. The baseline may depend on s, but it must not vary with a.
Two Quantities at One State
Fix a situation represented by s and consider several possible actions. The action value Q(s,a) is tied to whichever action a was selected, so it can have a different value for different actions. The baseline b(s) is tied to the situation rather than to the selected action. Therefore, for the same s, the baseline is the same reference across the actions being compared.
The diagram separates the roles of the two quantities. Each action has its own action value, while all of those action values are compared with the same b(s) for the fixed situation. The comparison can be written conceptually as Q(s,a) minus b(s).
The Baseline as a Reference
Comparing an Action with a State Reference
Suppose a fixed situation has an illustrative baseline of 6. Compare two illustrative action values: 9 for one action and 4 for another.
Choose the first action: The first action has an illustrative action value of 9. Relative to the baseline of 6, its comparison is 9 minus 6, which is 3.
Choose the second action: The second action has an illustrative action value of 4. Relative to the same baseline of 6, its comparison is 4 minus 6, which is negative 2.
Interpret the comparison: The first action is above the reference, while the second action is below it. The baseline provides the common point of comparison; it does not replace either action value.
The illustrative relative comparisons are 3 and negative 2. The numbers are examples of the comparison rule, not a prescribed baseline.
Why the Theorem Stays Valid
The generalized policy-gradient theorem uses the difference between the action value and the baseline in place of the action value by itself. The key condition is that b(s) must not vary with a. It may depend on the situation represented by s, and the source also allows it to be a random variable, but it must remain action-independent.
The reason the theorem remains valid is that the quantity introduced by subtracting an action-independent baseline is zero in the theorem's derivation. The action-independence condition is therefore not a minor preference. It is the condition that permits the baseline to be included without changing the validity of the theorem.
Tracing the Action Check
Two Candidate Baselines
For one fixed situation s, inspect two proposed references. Candidate A is b(s), and Candidate B changes when the selected action changes.
Inspect Candidate A: Candidate A is written as b(s). It is associated with the situation and does not vary with the action, so it satisfies the required condition.
Inspect Candidate B: Candidate B changes with the selected action. It is therefore action-dependent and does not satisfy the condition required for the generalized theorem.
Apply the theorem check: The correct question is not whether the reference has a particular numerical value. The correct question is whether it varies with a.
Candidate A qualifies under the stated condition; Candidate B does not.
| Quantity | Associated with | May vary across actions? | Role |
|---|---|---|---|
| Q(s,a) | Selected action | Yes | Action value |
| b(s) | Situation s | No | Reference for comparison |
| Action-dependent reference | Situation and action | Yes | Fails the required baseline condition |
The defining distinction is whether the quantity varies with the action.
When reading or applying the generalized theorem, explicitly check whether the proposed baseline depends on the action. Do not decide that a baseline is valid merely because it is called a baseline or because it has a convenient numerical value.
Common Mistakes
Treating the baseline as a replacement for the action value.
The generalized theorem compares the action value with the baseline; the baseline supplies the reference rather than replacing the action value.
Fix:
Keep both quantities in the comparison: the selected action's value and the action-independent reference.Assuming the baseline must be a fixed constant everywhere.
The baseline may depend on the relevant situation represented by s.
Fix:
Check whether it stays unchanged across actions for the same s. It does not have to be identical across all situations.Allowing a reference that changes with the selected action.
The required action-independence condition fails.
Fix:
Use a quantity that does not vary with a for the situation being considered.Thinking the baseline must be deterministic.
The source allows the baseline to be a random variable.
Fix:
The decisive condition is that it does not vary with the action, not that it must be a single fixed number.
Practice Check
For a fixed situation s, decide whether each proposed reference satisfies the baseline condition: (1) a quantity that depends only on s, (2) a random quantity that does not vary with the selected action, and (3) a quantity that changes whenever a changes. Explain why each one does or does not qualify.
Hints
- Focus on whether the reference varies with the action for the same situation.
- A baseline may depend on s and may be a random variable.
- The theorem's validity depends on action-independence.
- A qualifying baseline is an action-independent reference associated with the situation. It is combined with the action value through their difference, and the generalized policy-gradient theorem remains valid because the subtracted contribution is zero under the action-independence condition.
Key Takeaways
- The baseline b(s) extends the policy-gradient theorem by providing a reference for comparing an action value.
- Q(s,a) is associated with the selected action, while b(s) is associated with the situation.
- For a fixed situation, a valid baseline does not vary with the action.
- The baseline may depend on s and may even be a random variable, as long as it remains action-independent.
- The theorem remains valid because the contribution introduced by subtracting an action-independent baseline is zero in the derivation.