Concepts / State-Value Estimates

State-Value Estimates

Bootstrapping is the practice of revising an estimate on the basis of another estimate.

  • Programming

A Value Estimate Can Learn from Another Estimate

When estimating the value of a state, an update does not always have to wait for a final answer. The estimate for an earlier state can instead be revised using an estimate associated with a later state. This practice is called bootstrapping.

Bootstrapping means revising an estimate on the basis of another estimate.

Following the Estimate Backward

supplies informationis revisedState Aestimate being reconsideredState Bsuccessor estimateRevised State Aupdated estimate
How does the estimate for the current state change when it is updated using the estimated value of a successor state?

Imagine that State A is the state currently being reconsidered and State B is its successor. The estimate associated with State B supplies information used to revise the estimate for State A. The important point is not a particular numerical rule. The important point is the direction of dependence: one estimate becomes part of the basis for changing another estimate. That is bootstrapping.

Revising State A Using State B

Suppose an agent is estimating the value of State A and has an estimate for successor State B. How does this illustrate bootstrapping?

Identify the current estimate: State A is the state whose value estimate is being reconsidered.

Identify the supporting estimate: State B is a later state, and its value estimate supplies information for the update.

Describe the revision: The estimate for State A is revised using the estimate for State B rather than being revised directly from a final answer.

Name the mechanism: Because one estimate is updated on the basis of another estimate, the update is bootstrapping.

The estimate for State A is bootstrapped from the estimate for successor State B.

Dynamic Programming Moves Information from Successors

Dynamic Programming applies bootstrapping systematically. Its state-value updates use estimates of successor states. As a result, estimates farther along a state sequence provide information for updating the state currently being reconsidered. The information therefore moves from successor-state estimates toward the current state estimate.

feedsproducesSuccessor estimateState BState-value updateuses successor informationUpdated estimateState A
How do successor-state values flow backward to update the value estimate of the current state?

All Dynamic Programming methods share this update pattern: they update state-value estimates from estimates of successor states.

Model Dependence and Bootstrapping

Two separate questions help classify a reinforcement learning method. First, does it require a complete and accurate model of the environment? Second, does it bootstrap by updating one estimate from another? These questions describe different properties. Requiring a model and bootstrapping are not the same characteristic.

requiresusesdoes not requiredoes not usedoes not requireusesDynamic Programmingmodel required; bootstrapsComplete accuratemodelmodel requirementOther methodno model; no bootstrappingBootstrappingestimate update propertyOther methodno model; bootstraps
Which reinforcement-learning methods bootstrap, which require a complete and accurate environment model, and can these properties vary independently?
Method propertyWhat it asksSource-supported example
Complete and accurate modelDoes the method require a complete and accurate model of the environment?Dynamic Programming requires one
BootstrappingDoes the method revise one estimate using another estimate?Dynamic Programming does this
No model and no bootstrappingCan a method have neither property?Some reinforcement learning methods can
No model and bootstrappingCan a method bootstrap without requiring a model?Some reinforcement learning methods can

What the Environment Model Question Means

model informationsupportsEnvironmentcomplete and accurateValue-estimationmethoduses a modelState-value estimateupdated or evaluated
What environment information must be available when a value-estimation method requires a complete and accurate model?

A complete and accurate environment model is a requirement about the information available about the environment. It is separate from the question of whether an estimate is updated from another estimate. Dynamic Programming has both properties: it requires a complete and accurate model and it bootstraps. Other reinforcement learning methods may not require a model while still bootstrapping, or may have neither property.

Mistakes in Classification

  • Treating bootstrapping as waiting for a final answer

    Bootstrapping specifically means revising an estimate using another estimate, such as a successor-state estimate.

    Fix: Ask whether one estimate is being used as part of the basis for revising another estimate.

  • Assuming every use of a model is bootstrapping

    Model dependence and bootstrapping are separate properties.

    Fix: Check the model requirement and the estimate-update pattern as two independent questions.

  • Missing bootstrapping in Dynamic Programming

    Dynamic Programming also updates state-value estimates from estimates of successor states.

    Fix: Remember that Dynamic Programming has both properties: it requires a complete and accurate model and it bootstraps.

Classify the Method

EASY

A method revises the estimate for a current state using the estimate of a successor state. The method does not require a complete and accurate model of the environment. Which property does the method have, and which property does it not have?

Hints
  • Look separately at how the estimate is updated.
  • Then look separately at whether a complete and accurate model is required.

What do you think happens?

What should you conclude before checking the explanation?

  • It bootstraps but does not require a complete and accurate model
  • It requires a complete and accurate model but does not bootstrap
  • It has neither property
Reveal answer

Answer: It bootstraps but does not require a complete and accurate model.

Updating one estimate from a successor-state estimate is bootstrapping. The separate statement that no complete and accurate model is required determines the model property.

Key Takeaways

  1. Bootstrapping revises an estimate using another estimate.
  2. Dynamic Programming updates a state-value estimate using estimates of successor states.
  3. Dynamic Programming requires a complete and accurate environment model and also bootstraps.
  4. Requiring a model and bootstrapping are separate properties of reinforcement learning methods.
  5. A method can bootstrap without requiring a complete and accurate model.

Key Takeaways

  • Bootstrapping is the practice of revising an estimate on the basis of another estimate.
  • Dynamic Programming uses successor-state estimates to update state-value estimates.
  • A complete and accurate environment model is a separate requirement from bootstrapping.
  • Reinforcement learning methods can differ independently in whether they require a model and whether they bootstrap.