State-Value Estimates
Bootstrapping is the practice of revising an estimate on the basis of another estimate.
A Value Estimate Can Learn from Another Estimate
When estimating the value of a state, an update does not always have to wait for a final answer. The estimate for an earlier state can instead be revised using an estimate associated with a later state. This practice is called bootstrapping.
Bootstrapping means revising an estimate on the basis of another estimate.
Following the Estimate Backward
Imagine that State A is the state currently being reconsidered and State B is its successor. The estimate associated with State B supplies information used to revise the estimate for State A. The important point is not a particular numerical rule. The important point is the direction of dependence: one estimate becomes part of the basis for changing another estimate. That is bootstrapping.
Revising State A Using State B
Suppose an agent is estimating the value of State A and has an estimate for successor State B. How does this illustrate bootstrapping?
Identify the current estimate: State A is the state whose value estimate is being reconsidered.
Identify the supporting estimate: State B is a later state, and its value estimate supplies information for the update.
Describe the revision: The estimate for State A is revised using the estimate for State B rather than being revised directly from a final answer.
Name the mechanism: Because one estimate is updated on the basis of another estimate, the update is bootstrapping.
The estimate for State A is bootstrapped from the estimate for successor State B.
Dynamic Programming Moves Information from Successors
Dynamic Programming applies bootstrapping systematically. Its state-value updates use estimates of successor states. As a result, estimates farther along a state sequence provide information for updating the state currently being reconsidered. The information therefore moves from successor-state estimates toward the current state estimate.
All Dynamic Programming methods share this update pattern: they update state-value estimates from estimates of successor states.
Model Dependence and Bootstrapping
Two separate questions help classify a reinforcement learning method. First, does it require a complete and accurate model of the environment? Second, does it bootstrap by updating one estimate from another? These questions describe different properties. Requiring a model and bootstrapping are not the same characteristic.
| Method property | What it asks | Source-supported example |
|---|---|---|
| Complete and accurate model | Does the method require a complete and accurate model of the environment? | Dynamic Programming requires one |
| Bootstrapping | Does the method revise one estimate using another estimate? | Dynamic Programming does this |
| No model and no bootstrapping | Can a method have neither property? | Some reinforcement learning methods can |
| No model and bootstrapping | Can a method bootstrap without requiring a model? | Some reinforcement learning methods can |
What the Environment Model Question Means
A complete and accurate environment model is a requirement about the information available about the environment. It is separate from the question of whether an estimate is updated from another estimate. Dynamic Programming has both properties: it requires a complete and accurate model and it bootstraps. Other reinforcement learning methods may not require a model while still bootstrapping, or may have neither property.
Mistakes in Classification
Treating bootstrapping as waiting for a final answer
Bootstrapping specifically means revising an estimate using another estimate, such as a successor-state estimate.
Fix:
Ask whether one estimate is being used as part of the basis for revising another estimate.Assuming every use of a model is bootstrapping
Model dependence and bootstrapping are separate properties.
Fix:
Check the model requirement and the estimate-update pattern as two independent questions.Missing bootstrapping in Dynamic Programming
Dynamic Programming also updates state-value estimates from estimates of successor states.
Fix:
Remember that Dynamic Programming has both properties: it requires a complete and accurate model and it bootstraps.
Classify the Method
A method revises the estimate for a current state using the estimate of a successor state. The method does not require a complete and accurate model of the environment. Which property does the method have, and which property does it not have?
Hints
- Look separately at how the estimate is updated.
- Then look separately at whether a complete and accurate model is required.
What do you think happens?
What should you conclude before checking the explanation?
Reveal answer
Answer: It bootstraps but does not require a complete and accurate model.
Updating one estimate from a successor-state estimate is bootstrapping. The separate statement that no complete and accurate model is required determines the model property.
Key Takeaways
- Bootstrapping revises an estimate using another estimate.
- Dynamic Programming updates a state-value estimate using estimates of successor states.
- Dynamic Programming requires a complete and accurate environment model and also bootstraps.
- Requiring a model and bootstrapping are separate properties of reinforcement learning methods.
- A method can bootstrap without requiring a complete and accurate model.
Key Takeaways
- Bootstrapping is the practice of revising an estimate on the basis of another estimate.
- Dynamic Programming uses successor-state estimates to update state-value estimates.
- A complete and accurate environment model is a separate requirement from bootstrapping.
- Reinforcement learning methods can differ independently in whether they require a model and whether they bootstrap.