Concepts / Dynamic Programming Methods

Dynamic Programming Methods

Bootstrapping is the practice of revising an estimate on the basis of another estimate.

  • Programming

Why Successor Estimates Matter

Dynamic Programming methods revise a state-value estimate by using estimates associated with successor states. This means that value information can move from a later state toward the state currently being reconsidered. The central idea behind this process is bootstrapping: revising one estimate on the basis of another estimate.

Dynamic Programming methods use successor-state estimates when updating state-value estimates.

Bootstrapping a State Estimate

Bootstrapping is the practice of revising an estimate on the basis of another estimate.

Suppose State A is the state whose value estimate is being reconsidered. Instead of revising A directly from a final answer, the method looks at an estimate belonging to a later state, such as successor State B. Information from B becomes part of the basis for revising A. Because one estimate helps revise another estimate, the update bootstraps.

supplies estimatevalue updateState Ainitial estimateState Arevised estimateState Bsuccessor estimate
How does a current value estimate get revised using another estimate of a successor state?

State A and Its Successor

Explain why revising State A using an estimate for successor State B is bootstrapping.

Identify the current state: State A is the state whose value estimate is being reconsidered.

Identify the successor: State B is a later state and its value estimate supplies information for the update.

Apply the update: The estimate for A is revised using the estimate for B rather than being revised directly from a final answer.

Classify the update: Because one estimate is used to revise another estimate, the update is a bootstrapping update.

The update bootstraps the estimate for State A from the estimate for successor State B.

Information Moving Backward

Dynamic Programming applies bootstrapping systematically. Its state-value updates are based on estimates of successor states, so the estimate for one state is connected to estimates farther along the state sequence. During an update, information from successor states moves toward the state currently being reconsidered. This makes the updates internally connected: an estimate for one state can provide the basis for changing an estimate for another state.

estimateestimateproducesState Bsuccessor estimateState-value updatecurrent stateRevised stateestimateupdated valueState Csuccessor estimate
How does value information move from successor states back to the current state during an update?

The Environment Model

Dynamic Programming requires a complete and accurate model of the environment. That model is used to consider possible successor states while state-value estimates are updated. The model requirement describes what information the method needs about the environment; it is a separate question from whether the method bootstraps.

supports consideration ofhaveinformComplete accuratemodelenvironment informationSuccessor statespossible statesSuccessor estimatesvalue informationState-value updatecurrent state
What information must an environment model provide, and how is that model used to consider successor states?

Two Independent Questions

When classifying a reinforcement learning method, ask two separate questions. First, does it require a complete and accurate model of the environment? Second, does it bootstrap by updating one estimate from another estimate? Dynamic Programming answers yes to both questions, but these properties do not always appear together in other methods.

yesyesnononoyesDynamic Programmingmodel: yes; bootstrap: yesComplete accuratemodelrequired?Method Amodel: no; bootstrap: noEstimate fromestimateused?Method Bmodel: no; bootstrap: yes
How can reinforcement learning methods differ independently in whether they bootstrap and whether they require a complete, accurate environment model?
Method categoryRequires a complete and accurate modelBootstraps
Dynamic ProgrammingYesYes
Another reinforcement learning methodNoNo
Another reinforcement learning methodNoYes

Estimate or Observed Outcome

The defining distinction is what supplies the basis for an update. In bootstrapping, an estimate for another state supplies that basis. This differs from revising an estimate directly from a final answer or completed outcome. Dynamic Programming uses the first pattern: successor-state estimates support the update of the current state estimate.

basis forrevised asbasis forrevised asState A estimatebeing revisedSuccessor estimateanother estimateBootstrappingestimate from estimateFinal answercompleted outcomeDirect revisionfrom final answer
What is the difference between updating a value from another estimate and updating it from a completed return or observed outcome?

Check Your Classification

EASY

A method revises the value estimate for State A using an estimate for successor State B. The method also has a complete and accurate model of the environment. Which two properties does this description identify?

Hints
  • Look separately at the source of the update and the information the method requires.
  • Using one estimate to revise another is the definition of bootstrapping.

What do you think happens?

Does the method bootstrap, require a complete and accurate model, or both?

  • It bootstraps only
  • It requires a model only
  • It does both
  • Neither property is identified
Reveal answer

Answer: It does both.

The update uses a successor-state estimate, so it bootstraps. The description also explicitly gives the method a complete and accurate environment model.

Key Takeaways

  1. Bootstrapping means revising an estimate using another estimate.
  2. Dynamic Programming updates a state-value estimate using estimates of successor states.
  3. The update moves value information from successor states toward the state being reconsidered.
  4. Dynamic Programming requires a complete and accurate environment model and also bootstraps.
  5. Model dependence and bootstrapping are separate properties, so reinforcement learning methods can combine them in different ways.

Key Takeaways

  • Bootstrapping revises one estimate on the basis of another estimate.
  • Dynamic Programming uses successor-state estimates to update current state-value estimates.
  • Dynamic Programming requires a complete and accurate environment model.
  • Requiring a model and bootstrapping are independent classification questions.
  • A method may differ from Dynamic Programming by changing either property or both.