Concepts / Efficient Action-Value Estimation

Efficient Action-Value Estimation

An average can be updated incrementally rather than recomputed from all previous rewards.

  • Machine Learning

A New Reward Arrives

Suppose you are estimating the value of an action from the rewards it produces. When another reward arrives, you could recompute the average from every reward observed so far. That works, but it repeats earlier work. An incremental update takes a more efficient approach: it uses the existing estimate, the reward count, and only the new reward.

combinesets step sizenew informationExisting estimateQ_nUpdated estimatenew averageReward countnNew rewardR_n
What changes in the estimate when a new reward arrives, and how does this avoid revisiting every previous reward?

The Incremental Average

For an average, the old estimate is Q_n, the new reward is R_n, and the step size is 1/n. The reward count determines how large a step the update takes. Rather than storing and revisiting every previous reward, the method updates Q_n using the current estimate, the count, and the new reward.

Q_n ← Q_{n-1} + (1/n)[R_n - Q_{n-1}]

NewEstimateresult←assignmentOldEstimatestarting value+addStepSizeupdate amountTargetnew information-differenceOldEstimatestarting value
How does an old estimate become a new estimate by adding a step-sized portion of the target error?

Following the Adjustment

Moving an Estimate Toward a Reward

An old estimate is 4 and the new reward is 10. In the source example, the estimate moves toward 10 by an adjustment of 2.

Start with the old estimate: The current estimate is 4.

Compare with the target: The new reward is 10, so the target is above the old estimate.

Use the error: The difference between the reward and the old estimate determines the direction of the update.

Apply the adjustment: The estimate increases by 2 rather than jumping all the way to the reward.

The estimate changes from 4 to 6. It moves toward 10 by an adjustment of 2.

moves upwarddifference is scaledadded4old estimate6updated estimate10target reward2adjustment
How does the difference between the target reward and the old estimate determine the direction and size of the update?

From Average to General Rule

The average-specific update is one instance of a broader pattern: NewEstimate ← OldEstimate + StepSize [Target - OldEstimate]. The bracketed quantity is the error between what the estimate currently says and what the target says. The step size determines how much of that error is used to change the estimate.

comparetargetdeterminescale and addscalestarting pointQ_nexisting estimateTarget errorR_n - Q_nUpdated action valuenew estimateR_nnew rewardStep size1/nnreward count
How do the existing estimate, new reward, and reward count flow through the calculation to produce the updated action value?

When applying the rule, name the three important inputs before calculating: the existing estimate, the new reward or target, and the reward count. Then identify the error before applying the step size. This keeps the direction of the update visible.

Common Calculation Mistakes

  • Replacing the estimate with the new reward

    The incremental rule adds a step-sized portion of the error; it does not generally jump all the way to the target.

    Fix: Calculate the difference between the target and the old estimate, scale that difference, and add the adjustment to the old estimate.

  • Ignoring the error term

    The update depends on how far the target is from the old estimate.

    Fix: Form the target-minus-old-estimate difference first.

  • Recomputing every previous reward

    That repeats work that the incremental rule is designed to avoid.

    Fix: Keep the current estimate and reward count, then update them with the new reward.

  • Forgetting the role of the reward count

    For the average-specific rule, the step size is 1/n, so the count controls the scale of the update.

    Fix: Use the reward count to determine the step size before applying the error.

Try the Update

MEDIUM

An existing estimate is 6, the new reward is 10, and the reward count determines a step size of 1/4. Use the incremental pattern to find the adjustment and the updated estimate.

Hints
  • First identify the target error by subtracting the old estimate from the target.
  • Multiply that error by the step size.
  • Add the adjustment to the old estimate.
  1. Efficient action-value estimation keeps an existing estimate instead of recomputing an average from every previous reward. A new reward supplies a target, and the difference between that target and the old estimate is the error term. The estimate moves by a step-sized portion of that error. For an average, the step size is 1/n, and the needed state is the current estimate Q_n and the reward count n.

Key Takeaways

  • An average can be updated incrementally without revisiting every previous reward.
  • The update uses the old estimate, the new reward, and the reward count.
  • The error term is the target minus the old estimate.
  • The general pattern is NewEstimate ← OldEstimate + StepSize [Target - OldEstimate].
  • For an average, the step size is 1/n, so the reward count controls the size of the adjustment.