Value Functions in Reinforcement Learning
Evolutionary methods use policy search and reward comparison rather than value-function estimation.
Two Routes to Better Decisions
A reinforcement learning method needs a way to favor behavior that produces better results. One route estimates how valuable states or actions are. Another route, used by evolutionary methods, evaluates complete policies by observing the reward an agent obtains during its lifetime. These routes can address reinforcement learning problems without relying on the same unit of evaluation.
A Population of Policies
Imagine a collection of agents that do not improve through learning during their individual lifetimes. Each agent uses a different policy to decide how it interacts with the environment. The method observes the behavior produced by each policy over the agent's interaction period, records the reward obtained, and uses those results to favor stronger candidates.
- Create or consider a group of candidate policies.
- Let each candidate agent interact with the environment during its lifetime.
- Evaluate the total reward produced by each complete behavior.
- Compare candidates according to the reward they obtained.
- Favor stronger candidates as the basis for further policy search.
What Evolutionary Evaluation Measures
A value estimate is a long-run prediction of total future reward. Evolutionary evaluation instead assigns importance to the reward obtained by an agent's complete behavior over its interaction period.
| Evaluation viewpoint | What is evaluated | What the result represents |
|---|---|---|
| Evolutionary method | A complete policy and its lifetime behavior | The reward obtained over the interaction period |
| State-value estimate | A state | Predicted total future reward from that state |
| Action-value estimate | A particular action in a state | Predicted total future reward from taking that action |
State Values and Action Values
A state-value estimate, written as V, describes a state by predicting the total future reward expected from that state. An action-value estimate, written as Q, describes a particular action by predicting the total future reward expected from taking that action. Both are long-run forecasts rather than descriptions limited to the next immediate result.
Choosing Between Two Actions
Comparing State Values
An agent is considering two actions. One action leads to state A, whose value estimate is 8. The other leads to state B, whose value estimate is 3. Which action is preferred when the agent uses state values?
Identify the evaluated outcomes: The first action is evaluated through state A, and the second action is evaluated through state B.
Compare the estimates: State A has the larger value estimate: 8 is greater than 3.
Select the associated action: Using state values means favoring the action that leads to the highest-valued state.
The agent prefers the action that leads to state A. The estimates represent predicted total future rewards, not immediate rewards.
The comparison point changes depending on the estimate available. With V, the agent compares the values of states reached by the candidate actions. With Q, the agent compares the estimates attached directly to the actions. In either case, the preferred choice is the one associated with the larger predicted total future reward.
What do you think happens?
An agent has two possible actions. Under a state-value approach, the first leads to a state valued at 8 and the second leads to a state valued at 3. Which action is preferred?
Reveal answer
Answer: The action leading to the state valued at 8
State-value decision-making favors the action that leads to the highest-valued state. The values are long-run predictions of total future reward.
When Evolutionary Search Fits
Evolutionary methods may be effective when the policy space is small, when good policies are common or easy to find within that space, or when substantial time is available for searching. These conditions make it more plausible that comparison and selection will encounter useful policies.
They may also help when the learning agent cannot accurately sense the state of its environment. Methods that depend on accurate environmental sensing can face difficulty in that situation, while evolutionary evaluation can still compare the behavior produced by different policies according to the reward obtained.
Mistakes in Value Comparisons
Treating a value estimate as the immediate reward from the next step.
A value estimate is a prediction of total future reward over the long run, not a description limited to the next immediate result.
Fix:
Read the estimate as a forecast of the total reward available from the state or action in the future.Saying that V and Q evaluate exactly the same object.
V assigns an estimate to a state, while Q assigns an estimate to a specific action.
Fix:
Ask whether the estimate describes the situation itself or a particular choice in that situation.Assuming evolutionary methods must construct a value function.
Evolutionary methods compare complete policies through the rewards obtained by their lifetime behavior.
Fix:
Track the unit of evaluation: whole-policy lifetime reward for evolutionary search, or state/action forecasts for value-based evaluation.Comparing a state value directly with an action value as though they were the same kind of entry.
State values and action values place the comparison at different points: reached states for V and actions themselves for Q.
Fix:
Use V estimates to compare the states reached by choices, or use Q estimates to compare the choices directly.
Practice Check
An evolutionary method evaluates three candidate agents after their interaction periods. Candidate R obtains more reward than candidates S and T. Separately, an agent faces two actions: one leads to a state with V equal to 6, and the other leads to a state with V equal to 2. Explain which candidate is favored in the evolutionary process and which action is preferred under the state-value approach.
Hints
- For the population, focus on complete lifetime behavior and the reward obtained.
- For the two actions, compare the values of the states reached.
- Remember that the values predict total future reward rather than immediate reward.
Practice Solution
Use the population reward and the two state-value estimates to identify the preferred choices.
Evaluate the population: Candidate R is favored because it obtained more reward from its complete lifetime behavior.
Evaluate the state-value choice: The action leading to the state valued at 6 is preferred because 6 is greater than 2.
Keep the viewpoints separate: The first decision compares complete-agent outcomes, while the second compares values of states reached by actions.
Candidate R is favored in the evolutionary search, and the action leading to the state with V equal to 6 is preferred under the state-value approach.
Key Takeaways
- Evolutionary methods search over policies by comparing the rewards produced by complete lifetime behaviors.
- Selection favors candidates that obtain more reward, without requiring a value estimate for each state or action.
- A state-value estimate V predicts total future reward from a state.
- An action-value estimate Q predicts total future reward from taking a particular action in a state.
- With V, compare the values of states reached by actions; with Q, compare the action values directly.
Key Takeaways
- Evolutionary reinforcement learning evaluates whole policies through lifetime reward rather than estimating values for individual states or actions.
- Population-based search generates behavior, compares rewards, and favors stronger candidates for further policy search.
- V predicts total future reward from a state, while Q predicts total future reward from a particular action.
- State-value choice compares the states reached by actions; action-value choice compares estimates attached directly to actions.
- Evolutionary methods may be useful in small or searchable policy spaces, when substantial search time is available, or when accurate state sensing is difficult.