Concepts / Policies in Reinforcement Learning

Policies in Reinforcement Learning

A tabular method stores value-function estimates directly in arrays or tables.

  • Programming

Why Problem Size Matters

A reinforcement learning problem may involve states, actions, or combinations of both. If the collection is small, a method can store one value estimate for every relevant state, or for every relevant state and action combination. This is the basic idea behind a tabular method: represent value-function estimates directly in an array or table.

The important question is not whether a table is convenient to draw. The question is whether the state and action spaces are small enough that the complete collection of relevant estimates can be represented explicitly. When that condition holds, tabular methods can often determine the optimal value function and the optimal policy exactly. When the problem is much larger, approximate methods are used instead; they handle larger problems but provide approximate solutions.

supportsmotivatesSmall state andaction spacesvalues stored directlyArray or tableexplicit value estimatesMuch largerproblemsfull table impracticalApproximate methodapproximate solution
How does problem size affect the choice between an explicit table and an approximate representation?

Mapping States to Values

A value function represents estimates associated with states or actions. In a tabular representation, each relevant state can be associated with a particular position in an array or with an entry in a table. The content of that entry is the current value estimate for that state. If the method represents state-action values, the entries instead correspond to relevant combinations of a state and an action.

A Generated Small-State Table

Consider a generated toy problem with three relevant states: Start, Choice, and Goal. How can a tabular method represent its current value estimates?

Identify the entries: Because the generated problem has only a handful of relevant states, the method can reserve one value entry for Start, one for Choice, and one for Goal.

Store estimates: The entries hold the current estimates associated with those states. They are estimates used by the method, not permanent labels attached to the states.

Use the representation: As the method learns, its updates change the entries. The collection of entries is the direct tabular representation of the value function.

A small state collection makes it practical to store the value-function estimates explicitly in an array or table.

maps tomaps tomaps toStartstateValue entryStart estimateChoicestateValue entryChoice estimateGoalstateValue entryGoal estimate
How does each state map to a specific position containing its current value estimate?

Exact and Approximate Solutions

Exact solutionApproximate solution
Precisely optimal value function and policy when the tabular approach successfully appliesValue function and policy are estimates rather than a precisely optimal solution
Associated with problems small enough for explicit value representationUsed for much larger problems
The stored solution is not merely an approximation of the optimumThe method handles scale by providing an approximate solution

Exact does not mean that the table avoids learning or solving. It means that, when the tabular approach applies successfully, the resulting value function and policy are precisely optimal rather than merely approximate. A table is useful because the problem is small enough to represent explicitly; it does not remove the need to solve the reinforcement learning problem.

can producecan produceusesusesSmall problemexplicit representationOptimal valuefunctionprecisely optimalMuch larger problemapproximate representationApproximate valuefunctionestimateOptimal policyprecisely optimalApproximate policyestimate
What differs between storing the precisely optimal solution and storing estimates that may contain approximation?

Backing Up Along Trajectories

A trajectory supplies an ordered path through states. A backup carries value information along that path so that estimates associated with states can be updated. The path may be an actual trajectory produced while the method interacts with a problem, or a possible trajectory considered by the method. The essential idea is movement of value information through an ordered sequence, not a particular update formula.

For example, suppose a generated trajectory visits State A, then State B, then State C. Information associated with a later part of that path can be backed up toward earlier states. The estimate for State B can be changed using information from the later path, and information can also reach State A through the same trajectory. The source describes this as changing existing value estimates along actual or possible state trajectories.

backupbackupState Aearlier estimateState Bintermediate estimateState Clater value information
How does value information from a later state move backward through a sequence to update earlier value estimates?

Generalized Policy Iteration

Generalized policy iteration, or GPI, is the broad strategy that connects value estimation with behavior. A method maintains an approximate value function and an approximate policy. It repeatedly tries to improve each one using the other: the value function helps assess the policy, while the policy gives the value estimates a behavioral direction.

The word generalized matters. GPI is a shared organizing strategy, not one narrowly specified procedure with identical update steps in every method. The common pattern is to maintain both approximate objects and use their relationship to seek improvement. Value updates can change the value function, and the revised value estimates can contribute to improving the policy. The improved policy then remains connected to further value estimation.

guidessupplieschangeshelps improveApproximate policybehavioral directionState trajectoryactual or possibleValue backupupdates estimatesApproximate valuefunctionassesses behavior
How do an approximate policy and an approximate value function repeatedly improve one another?

Following the Improvement Loop

In a generated small problem, a method has an approximate policy and an approximate value function. How do the two objects participate in generalized policy iteration?

Begin with behavior: The approximate policy supplies a current behavioral direction, which is connected to trajectories through states.

Estimate and update: The method uses value estimation and backups along actual or possible trajectories to change its approximate value function.

Use the revised values: The revised value estimates help assess the current behavior and contribute to improving the approximate policy.

Continue the relationship: The improved policy remains connected to further value estimation, so the two approximate objects continue to influence one another.

GPI is a continuing mutual-improvement pattern linking an approximate policy with an approximate value function.

Three Organizing Ideas

Organizing ideaQuestion it answersRole in the overall method
Value estimationWhat is being learned?Estimates the value of states or actions
Value backupsHow does information move?Carries value information along actual or possible state trajectories and updates estimates
Generalized policy iterationHow do estimates and behavior improve together?Links an approximate value function and an approximate policy in a continuing improvement process

These three ideas provide a way to compare reinforcement learning methods that may look different when studied individually.

The three ideas form one connected picture. Value functions describe what is being learned. Backups describe how estimates are changed along trajectories. GPI describes how the estimates and policy are developed together. Keeping these roles distinct prevents a common confusion: a value function is the learned estimate, a backup is the process that changes an estimate, and GPI is the broader improvement relationship between values and behavior.

producesupdatesimprovesdevelops withinformsparticipates inValue estimationwhat is learnedApproximate policybehaviorValue backupshow estimates changeApproximate valuefunctionstate or action estimatesGeneralized policyiterationhow values and behaviorimprove
How are value estimation, value backups, and generalized policy iteration connected?

Mistakes to Avoid

  • Assuming every reinforcement learning problem should use a table

    Tabular methods are intended for problems with small state and action spaces. Approximate methods are used for much larger problems.

    Fix: Assess the size of the state and action spaces before deciding whether explicit value entries are practical.

  • Thinking that a table automatically solves the problem

    The table only provides a direct representation for value-function estimates. The reinforcement learning problem still has to be solved.

    Fix: Treat the table as a representation, then consider how values will be estimated, backed up, and connected to policy improvement.

  • Confusing an estimate with a permanent property of a state

    A value is an estimate used by a reinforcement learning method, and backups update existing estimates.

    Fix: Remember that value estimates can change as information is backed up along trajectories.

  • Treating GPI as one fixed update procedure

    GPI is a shared organizing strategy that links an approximate policy and an approximate value function; it does not require identical update steps.

    Fix: Look for the mutual-improvement pattern between the policy and value function.

Check Your Understanding

MEDIUM

A generated reinforcement learning problem has a small collection of relevant states and actions. Explain why a tabular method may be appropriate, what its table represents, and how value backups and generalized policy iteration fit into the method.

Hints
  • Start with the size of the state and action spaces.
  • Describe the entries as value-function estimates associated with states or actions.
  • Explain that backups update estimates along actual or possible trajectories.
  • Explain that GPI links an approximate value function and an approximate policy so they can improve one another.

What do you think happens?

A value estimate for a later state changes during a trajectory. What is the central purpose of backing up that information?

  • To update estimates associated with states earlier on the trajectory
  • To permanently label the later state
  • To make the table unnecessary
  • To guarantee that every problem has an exact solution
Reveal answer

Answer: To update estimates associated with states earlier on the trajectory

A backup carries value information along an actual or possible state trajectory so that existing estimates can be updated. It does not by itself make every problem tabular or guarantee an exact solution.

Key Takeaways

  • Tabular methods store value-function estimates directly in arrays or tables and are practical when state and action spaces are small.
  • When a tabular approach applies successfully, it can often produce the precisely optimal value function and policy; approximate methods address much larger problems with approximate solutions.
  • Value backups carry information along actual or possible state trajectories to update existing value estimates.
  • Generalized policy iteration links an approximate value function with an approximate policy so that each can help improve the other.
  • Value estimation, value backups, and generalized policy iteration distinguish what is learned, how estimates change, and how behavior and estimates develop together.