Efficient Action-Value Estimation
An average can be updated incrementally rather than recomputed from all previous rewards.
A New Reward Arrives
Suppose you are estimating the value of an action from the rewards it produces. When another reward arrives, you could recompute the average from every reward observed so far. That works, but it repeats earlier work. An incremental update takes a more efficient approach: it uses the existing estimate, the reward count, and only the new reward.
The Incremental Average
For an average, the old estimate is Q_n, the new reward is R_n, and the step size is 1/n. The reward count determines how large a step the update takes. Rather than storing and revisiting every previous reward, the method updates Q_n using the current estimate, the count, and the new reward.
Q_n ← Q_{n-1} + (1/n)[R_n - Q_{n-1}]
Following the Adjustment
Moving an Estimate Toward a Reward
An old estimate is 4 and the new reward is 10. In the source example, the estimate moves toward 10 by an adjustment of 2.
Start with the old estimate: The current estimate is 4.
Compare with the target: The new reward is 10, so the target is above the old estimate.
Use the error: The difference between the reward and the old estimate determines the direction of the update.
Apply the adjustment: The estimate increases by 2 rather than jumping all the way to the reward.
The estimate changes from 4 to 6. It moves toward 10 by an adjustment of 2.
From Average to General Rule
The average-specific update is one instance of a broader pattern: NewEstimate ← OldEstimate + StepSize [Target - OldEstimate]. The bracketed quantity is the error between what the estimate currently says and what the target says. The step size determines how much of that error is used to change the estimate.
When applying the rule, name the three important inputs before calculating: the existing estimate, the new reward or target, and the reward count. Then identify the error before applying the step size. This keeps the direction of the update visible.
Common Calculation Mistakes
Replacing the estimate with the new reward
The incremental rule adds a step-sized portion of the error; it does not generally jump all the way to the target.
Fix:
Calculate the difference between the target and the old estimate, scale that difference, and add the adjustment to the old estimate.Ignoring the error term
The update depends on how far the target is from the old estimate.
Fix:
Form the target-minus-old-estimate difference first.Recomputing every previous reward
That repeats work that the incremental rule is designed to avoid.
Fix:
Keep the current estimate and reward count, then update them with the new reward.Forgetting the role of the reward count
For the average-specific rule, the step size is 1/n, so the count controls the scale of the update.
Fix:
Use the reward count to determine the step size before applying the error.
Try the Update
An existing estimate is 6, the new reward is 10, and the reward count determines a step size of 1/4. Use the incremental pattern to find the adjustment and the updated estimate.
Hints
- First identify the target error by subtracting the old estimate from the target.
- Multiply that error by the step size.
- Add the adjustment to the old estimate.
- Efficient action-value estimation keeps an existing estimate instead of recomputing an average from every previous reward. A new reward supplies a target, and the difference between that target and the old estimate is the error term. The estimate moves by a step-sized portion of that error. For an average, the step size is 1/n, and the needed state is the current estimate Q_n and the reward count n.
Key Takeaways
- An average can be updated incrementally without revisiting every previous reward.
- The update uses the old estimate, the new reward, and the reward count.
- The error term is the target minus the old estimate.
- The general pattern is NewEstimate ← OldEstimate + StepSize [Target - OldEstimate].
- For an average, the step size is 1/n, so the reward count controls the size of the adjustment.