Action-Value Methods
A sample average is a natural estimate of an action value.
From Rewards to an Estimate
When an action is selected several times, each selection produces a reward. A natural estimate of that action's value is the average of the rewards observed so far. The estimate is not based on one reward alone; it summarizes the experience collected from repeated selections.
Following the Estimate
A changing action-value estimate
An action produces the following observed rewards in sequence: 4, 6, and 2. Track the average estimate after each observation.
First observation: After observing only the reward 4, the average of the observed rewards is 4.
Second observation: After observing 4 and 6, the estimate is the average of those two rewards, which is 5.
Third observation: After observing 4, 6, and 2, the estimate is the average of all three rewards, which is 4.
The estimate changes whenever a new reward is included: 4, then 5, then 4.
This trace shows the central behavior of an action-value estimate: every new observation can change the average. The estimate after a particular observation represents the rewards observed up to that point, rather than rewards that may be collected later.
Reading the Notation
The source notation uses R_i for the reward received after the ith selection of the action. It uses Q_n for the estimated action value after the action has been selected n minus 1 times. Therefore, the rewards used for Q_n are the rewards observed before the estimate receives that label. The notation emphasizes that an estimate is tied to the experience available at a particular point in time.
When reasoning about an estimate, always identify which observations it includes. Confusing the number of selections with the number of rewards already included can lead to an incorrect interpretation of the estimate's label.
The Cost of Full History
The straightforward approach is to keep every reward and repeatedly add the entire collection whenever the average is needed. This approach has two growing costs. First, memory is needed to store the rewards. Second, computation is needed to add the rewards in the history. As experience accumulates, both the stored collection and the repeated addition become larger.
Incremental Computation
Incremental computation updates an estimate as new observations arrive instead of repeatedly treating the entire reward history as a fresh input. For action-value estimation, the desired behavior is constant memory and per-time-step computation. Constant memory means that the memory requirement does not grow simply because more rewards have been observed. Per-time-step computation means that processing each new time step remains bounded rather than requiring an ever-longer recomputation.
The key contrast is not whether the estimate represents the observed rewards. Both approaches can represent the same average. The contrast is whether each update revisits the entire history or maintains the estimate incrementally as observations arrive.
Common Reasoning Mistakes
Treating memory growth and computation growth as the same cost
The source identifies two separate growing costs: memory for storing rewards and computation for adding them.
Fix:
Analyze storage and arithmetic work separately.Assuming that an average is automatically maintained efficiently
The average is a natural estimate, but the method used to maintain it determines whether storage and work grow with experience.
Fix:
Distinguish the estimate itself from the procedure used to update it.Confusing constant memory with constant information about the past
Constant memory describes the desired behavior of the memory requirement, not a claim that earlier observations are irrelevant.
Fix:
Understand constant memory as avoiding memory growth simply because more rewards have been observed.
For every averaging method, ask two questions: What information must be stored, and how much work is required when the next reward arrives? These questions expose the difference between a straightforward full-history method and an incremental method.
Check Your Understanding
Suppose an action has produced several rewards. Describe, in your own words, what the action-value estimate represents, what grows when every reward is stored, and what constant memory and per-time-step computation are intended to prevent.
Hints
- Start with the average of the rewards observed so far.
- Separate the amount of stored reward history from the work of adding that history.
- Explain constant memory and per-time-step computation as goals for processing new observations.
What do you think happens?
If a new reward arrives, which approach requires revisiting the entire reward history: straightforward full-history averaging or incremental computation?
Reveal answer
Answer: Straightforward full-history averaging
The straightforward method repeatedly adds the entire collection of past rewards. Incremental computation updates the estimate as new observations arrive and aims for bounded per-time-step computation.
Key Takeaways
- A natural estimate of an action value is the average of the rewards observed after selecting that action.
- Keeping every reward and repeatedly adding the full history causes both memory and computation to grow as experience accumulates.
- Memory growth concerns how much reward data must be stored; computation growth concerns how much work is required to add that data.
- Incremental computation updates the estimate as new observations arrive instead of repeatedly treating the full history as a fresh input.
- Constant memory and per-time-step computation are desirable goals because they avoid costs that grow simply from collecting more rewards.
Key Takeaways
- An action value can be estimated by averaging the rewards observed from repeated selections of the action.
- Storing every reward creates growing memory use, while repeatedly adding the history creates growing computation.
- Incremental computation focuses on updating the estimate as each new observation arrives.
- Constant memory and per-time-step computation are the desired efficiency properties for action-value estimation.