Concepts / Value Estimation in Reinforcement Learning

Value Estimation in Reinforcement Learning

A reward signal describes immediate desirability, while a value function describes long-term desirability.

  • Programming

The tempting immediate choice

Imagine an agent choosing between two situations. One situation provides something desirable immediately. The other appears less attractive at first but tends to lead to better outcomes later. If the agent considers only the next moment, it chooses the first situation. Reinforcement learning needs a way to represent the longer-term consequence of beginning from each situation. That is the role of a value function.

chooseleads tochooseleads toDecisiontwo situationsSituation Alarge immediate rewardFuture outcomesless rewarding laterSituation Bsmall immediate rewardFuture outcomesmore rewarding later
How can a state with a small immediate reward lead to greater total future reward than a state with a large immediate reward?

Reward and value are different

A reward signal describes immediate desirability. It is supplied directly by the environment and represents what the agent receives on the current transition. A value estimate describes long-term desirability. It is a long-run prediction of the total future reward the agent can accumulate from a state or from taking a particular action.

different timescaleRewardcurrent transitionValue estimatefuture rewards
What is the difference between the reward received on the current transition and the accumulated future desirability predicted by a value estimate?

Reading a state value

The value of a state estimates the total reward the agent can expect to accumulate in the future from that state. It answers a forecasting question: if the agent is in this situation, how much future reward does it expect to receive?

leads torevealssupports predictionCurrent stateFuture observationsFuture rewardsState valuepredicted total futurereward
Starting from this state, what future rewards does the agent expect to receive over time?

The estimate is not limited to the reward visible right now. It represents the future reward expected from the state. This is why an agent can prefer a situation that is initially less attractive: the situation may be more promising when its later consequences are considered.

Reading an action value

V assigns an estimate to a state, while Q assigns an estimate to a specific action. An action value describes the total future reward expected from taking that particular action. Both are long-run predictions rather than descriptions limited to the next immediate result.

availableavailableleads toleads toCurrent stateAction AQ estimateFuture rewardsoutcomes after AAction BQ estimateFuture rewardsoutcomes after B
How do different actions from the same state lead to different future reward sequences and value estimates?

The distinction is about what is being evaluated. V asks how promising it is for the agent to be in a situation. Q asks how promising it is to take a particular action. The first estimate is attached to a state; the second is attached to an action.

Comparing decision routes

Choosing through state values

An agent is considering two actions. One action leads to state A, whose value estimate is 8. The other leads to state B, whose value estimate is 3. Which action is preferred when the agent uses state value estimates?

Identify the compared quantities: The agent compares the values of the states reached by the available actions, not the immediate rewards of those actions.

Compare the estimates: State A has an estimate of 8, while state B has an estimate of 3.

Select the larger estimate: The action leading to state A is preferred because 8 is the larger prediction of total future reward.

The agent prefers the action that leads to state A. The numbers 8 and 3 are predictions of future total reward, not immediate rewards.

comparecompareselect larger valueselect larger QState-value routecompare resulting statesStates A and B8 versus 3Preferred choicelarger relevant estimateAction-value routecompare available actionsActions A and BQ estimates
How does the agent select a decision when comparing the values of resulting states versus the values of available actions?
Decision viewpointWhat is estimatedHow the choice is guided
State valuesThe desirability of a stateFavor the action that leads to the highest-valued state
Action valuesThe desirability of a specific actionFavor the action with the highest action-value estimate

Why estimation is difficult

Rewards come directly from the environment, so identifying an immediate reward signal is comparatively straightforward. Values are different because they must be inferred from sequences of observations gathered across the agent's lifetime. The agent must repeatedly estimate and re-estimate what a state is worth as it encounters more observations and learns about the outcomes that tend to follow.

observe over timerevealssupportsCurrent stateObservation sequenceFollowing outcomesValue estimateexpected total futurereward
How does information about the current state and possible future transitions become a prediction of expected total reward?

Value estimation is central because an agent rarely knows the complete result of a decision immediately. Value estimates provide a forecast of which available choice appears to lead to the better long-term outcome.

Check your reasoning

EASY

An agent can choose between two situations. Situation A gives a larger reward immediately but tends to lead to less rewarding future states. Situation B gives a smaller reward immediately but tends to lead to more rewarding future states. Which situation may have the higher value, and why?

Hints
  • Separate the current reward from the prediction about future rewards.
  • Ask which situation tends to lead to the greater total future reward.

What do you think happens?

The action leading to state A has a state value estimate of 8. The action leading to state B has a state value estimate of 3. Which action is preferred when choosing through state values?

  • The action leading to state A
  • The action leading to state B
  • Neither action, because values describe only immediate rewards
Reveal answer

Answer: The action leading to state A.

When using state values, the agent favors the action that leads to the highest-valued state. The estimate of 8 is larger than 3, and both numbers predict total future reward rather than only the next reward.

Mistakes to avoid

  • Treating a value estimate as the immediate reward.

    The estimate is a prediction of total future reward, while the reward is supplied directly by the environment as an immediate signal.

    Fix: Describe the reward as immediate and the value as a long-run prediction.

  • Assuming that a low immediate reward means the state has low value.

    That state may tend to lead to future states offering greater rewards.

    Fix: Consider the total future reward predicted from the state.

  • Confusing a state value with an action value.

    V describes a state, whereas Q describes a specific action.

    Fix: Use V for the value of a state and Q for the value of taking a particular action.

  • Comparing the wrong objects when selecting an action.

    The state-value route compares the values of states reached by actions; the action-value route compares estimates attached directly to the actions.

    Fix: State clearly whether the decision is based on resulting-state values or available-action values.

Key takeaways

  1. A reward signal describes immediate desirability, while a value function describes long-term desirability.
  2. A state value estimates the total future reward expected from a state.
  3. An action value estimates the total future reward expected from taking a specific action.
  4. A state with a low immediate reward can still have high value when it tends to lead to more rewarding future states.
  5. Value estimates must be inferred from observation sequences and repeatedly updated as the agent learns about later outcomes.

Key Takeaways

  • Immediate rewards describe what the environment supplies now; value estimates predict total future reward.
  • State values evaluate situations, while action values evaluate particular choices.
  • An agent can prefer a lower immediate reward when the associated state or action promises better future outcomes.
  • Value estimation is challenging because it must be learned from sequences of observations and updated over time.