Concepts / Value Functions in Reinforcement Learning

Value Functions in Reinforcement Learning

Evolutionary methods use policy search and reward comparison rather than value-function estimation.

  • Programming

Two Routes to Better Decisions

A reinforcement learning method needs a way to favor behavior that produces better results. One route estimates how valuable states or actions are. Another route, used by evolutionary methods, evaluates complete policies by observing the reward an agent obtains during its lifetime. These routes can address reinforcement learning problems without relying on the same unit of evaluation.

producescompared byestimatesguidesEvolutionary methodpolicy searchLifetime behaviortotal rewardSelectionbetter candidatesValue-based methodvalue estimationStates or actionsestimated valueAction choicehigher estimate
How does an evolutionary method improve policies by evaluating complete-agent rewards instead of estimating values for states or actions?

A Population of Policies

Imagine a collection of agents that do not improve through learning during their individual lifetimes. Each agent uses a different policy to decide how it interacts with the environment. The method observes the behavior produced by each policy over the agent's interaction period, records the reward obtained, and uses those results to favor stronger candidates.

uses policiesreceivesis comparedfavorscontinues searchCandidate policiesagent groupAgent behaviorinteraction periodLifetime rewardbehavior resultReward comparisonstronger candidatesFurther policy searchselected basis
What happens as a group of agents generates behavior, receives lifetime rewards, gets compared, and produces the next generation of policies?
  1. Create or consider a group of candidate policies.
  2. Let each candidate agent interact with the environment during its lifetime.
  3. Evaluate the total reward produced by each complete behavior.
  4. Compare candidates according to the reward they obtained.
  5. Favor stronger candidates as the basis for further policy search.

What Evolutionary Evaluation Measures

A value estimate is a long-run prediction of total future reward. Evolutionary evaluation instead assigns importance to the reward obtained by an agent's complete behavior over its interaction period.

Evaluation viewpointWhat is evaluatedWhat the result represents
Evolutionary methodA complete policy and its lifetime behaviorThe reward obtained over the interaction period
State-value estimateA statePredicted total future reward from that state
Action-value estimateA particular action in a statePredicted total future reward from taking that action
producespredicts frompredicts fromPolicycomplete behaviorLifetime rewardone evaluationStatecurrent situationFuture rewardlong-run predictionActionparticular choice
What is the difference between assigning one total reward to an agent's complete lifetime behavior and predicting future returns from an individual state or action?

State Values and Action Values

A state-value estimate, written as V, describes a state by predicting the total future reward expected from that state. An action-value estimate, written as Q, describes a particular action by predicting the total future reward expected from taking that action. Both are long-run forecasts rather than descriptions limited to the next immediate result.

evaluated underhelps determinepredictsCurrent stateagent situationV estimatestate valueFuture rewardpredicted totalPolicybehavior rule
What future reward does a state-value estimate predict, and how does that prediction depend on the agent's current state and policy?
context forevaluated aspredictsStateagent situationQ estimateaction valueFuture rewardpredicted totalActionparticular choice
What future reward does an action-value estimate predict for taking a particular action in a particular state?

Choosing Between Two Actions

Comparing State Values

An agent is considering two actions. One action leads to state A, whose value estimate is 8. The other leads to state B, whose value estimate is 3. Which action is preferred when the agent uses state values?

Identify the evaluated outcomes: The first action is evaluated through state A, and the second action is evaluated through state B.

Compare the estimates: State A has the larger value estimate: 8 is greater than 3.

Select the associated action: Using state values means favoring the action that leads to the highest-valued state.

The agent prefers the action that leads to state A. The estimates represent predicted total future rewards, not immediate rewards.

leads toleads to8 is larger than 3compared with 8compare Qcompare QAction oneleads to state AState AV = 8Action oneQ estimatePreferred actionlarger relevant estimateAction twoleads to state BState BV = 3Action twoQ estimate
How does an agent use values to choose between actions when it has values for states versus values for state-action pairs?

The comparison point changes depending on the estimate available. With V, the agent compares the values of states reached by the candidate actions. With Q, the agent compares the estimates attached directly to the actions. In either case, the preferred choice is the one associated with the larger predicted total future reward.

What do you think happens?

An agent has two possible actions. Under a state-value approach, the first leads to a state valued at 8 and the second leads to a state valued at 3. Which action is preferred?

  • The action leading to the state valued at 8
  • The action leading to the state valued at 3
  • Neither action can be compared
Reveal answer

Answer: The action leading to the state valued at 8

State-value decision-making favors the action that leads to the highest-valued state. The values are long-run predictions of total future reward.

When Evolutionary Search Fits

Evolutionary methods may be effective when the policy space is small, when good policies are common or easy to find within that space, or when substantial time is available for searching. These conditions make it more plausible that comparison and selection will encounter useful policies.

They may also help when the learning agent cannot accurately sense the state of its environment. Methods that depend on accurate environmental sensing can face difficulty in that situation, while evolutionary evaluation can still compare the behavior produced by different policies according to the reward obtained.

Mistakes in Value Comparisons

  • Treating a value estimate as the immediate reward from the next step.

    A value estimate is a prediction of total future reward over the long run, not a description limited to the next immediate result.

    Fix: Read the estimate as a forecast of the total reward available from the state or action in the future.

  • Saying that V and Q evaluate exactly the same object.

    V assigns an estimate to a state, while Q assigns an estimate to a specific action.

    Fix: Ask whether the estimate describes the situation itself or a particular choice in that situation.

  • Assuming evolutionary methods must construct a value function.

    Evolutionary methods compare complete policies through the rewards obtained by their lifetime behavior.

    Fix: Track the unit of evaluation: whole-policy lifetime reward for evolutionary search, or state/action forecasts for value-based evaluation.

  • Comparing a state value directly with an action value as though they were the same kind of entry.

    State values and action values place the comparison at different points: reached states for V and actions themselves for Q.

    Fix: Use V estimates to compare the states reached by choices, or use Q estimates to compare the choices directly.

Practice Check

EASY

An evolutionary method evaluates three candidate agents after their interaction periods. Candidate R obtains more reward than candidates S and T. Separately, an agent faces two actions: one leads to a state with V equal to 6, and the other leads to a state with V equal to 2. Explain which candidate is favored in the evolutionary process and which action is preferred under the state-value approach.

Hints
  • For the population, focus on complete lifetime behavior and the reward obtained.
  • For the two actions, compare the values of the states reached.
  • Remember that the values predict total future reward rather than immediate reward.

Practice Solution

Use the population reward and the two state-value estimates to identify the preferred choices.

Evaluate the population: Candidate R is favored because it obtained more reward from its complete lifetime behavior.

Evaluate the state-value choice: The action leading to the state valued at 6 is preferred because 6 is greater than 2.

Keep the viewpoints separate: The first decision compares complete-agent outcomes, while the second compares values of states reached by actions.

Candidate R is favored in the evolutionary search, and the action leading to the state with V equal to 6 is preferred under the state-value approach.

Key Takeaways

  1. Evolutionary methods search over policies by comparing the rewards produced by complete lifetime behaviors.
  2. Selection favors candidates that obtain more reward, without requiring a value estimate for each state or action.
  3. A state-value estimate V predicts total future reward from a state.
  4. An action-value estimate Q predicts total future reward from taking a particular action in a state.
  5. With V, compare the values of states reached by actions; with Q, compare the action values directly.

Key Takeaways

  • Evolutionary reinforcement learning evaluates whole policies through lifetime reward rather than estimating values for individual states or actions.
  • Population-based search generates behavior, compares rewards, and favors stronger candidates for further policy search.
  • V predicts total future reward from a state, while Q predicts total future reward from a particular action.
  • State-value choice compares the states reached by actions; action-value choice compares estimates attached directly to actions.
  • Evolutionary methods may be useful in small or searchable policy spaces, when substantial search time is available, or when accurate state sensing is difficult.