Concepts / Value Estimates in Reinforcement Learning

Value Estimates in Reinforcement Learning

q∗(a) denotes the true value of action a.

  • Programming

Why Estimates Guide Decisions

An agent rarely knows the complete result of a decision immediately. Instead, it can treat each available choice as a forecast: which choice appears likely to produce the better long-term outcome? A value estimate provides that forecast by predicting the total reward the agent can accumulate in the future.

considerevaluatepredictsCurrent stateActionValue estimatePredicted total futurerewardFuture rewards
How does a value estimate connect a current situation or action to rewards received later?

True Values and Learning Estimates

In a bandit problem, two closely related quantities describe an action. The true value of action a is written q∗(a). An estimate of that true value at time t is written Qt(a). The two expressions concern the same action, identified by a, but they are not the same quantity. q∗(a) is the value being estimated, whereas Qt(a) is the agent's estimate at a particular time.

is estimated byq∗(a)True value of action aQt(a)Estimate at time t
How does the fixed true value q∗(a) differ from the agent's changing estimate Qt(a) at time t?
qaction valueQaction value estimate∗true valuettime of estimateaactionaaction
What does each symbol in q∗(a) and Qt(a) identify?

What Time Adds to an Estimate

The subscript t in Qt(a) tells you when the estimate is being considered. It does not identify a different action; the action is still identified by a. It identifies the time of the estimate. As the agent gains experience, its estimate can be considered at a later time, so the notation distinguishes the estimate at one time from an estimate at another time.

more experienceQt(a)estimate at time tQt+1(a)estimate at a later time
What changes in an action's estimate as the agent gains experience from time t to a later time?

What do you think happens?

An expression changes from Qt(a) to an estimate at a later time. What does the changed time notation tell you?

  • The action a has necessarily changed
  • The estimate is associated with a later time
  • The true value q∗(a) has been replaced
  • The notation now describes a state instead of an action
Reveal answer

Answer: The estimate is associated with a later time.

The subscript identifies the time of an estimate. The symbol a continues to identify the action, while q∗(a) remains the true value being estimated.

States, Actions, and Future Reward

A value estimate is a long-run prediction of total future reward. The distinction between state and action values is about what is being evaluated. V assigns an estimate to a state: it describes how promising it is for the agent to be in that situation. Q assigns an estimate to a specific action: it describes how promising it is to take that particular action. Neither description is limited to the next immediate result.

predicts from a statepredicts from an actionVvalue estimate for a stateQvalue estimate for anactionFuture rewardsFuture rewards
Does the estimate predict the value of being in a state generally, or the value of taking a particular action from that state?
EstimateWhat it evaluatesWhat the prediction concerns
VA stateTotal future reward expected from being in that state
QA specific actionTotal future reward expected from taking that action

Two Ways to Prefer a Decision

Comparing Two Decisions

An agent is considering two actions. One action leads to state A, whose state value estimate is 8. The other action leads to state B, whose state value estimate is 3. Which action is preferred when the agent uses state value estimates?

Identify the relevant estimates: The first decision leads to state A with estimate 8. The second decision leads to state B with estimate 3.

Compare the predictions: The estimates represent predicted total future reward, not merely immediate rewards. Since 8 is larger than 3, state A has the larger predicted total future reward.

Select the preferred decision: Using state value estimates, the agent prefers the action that leads to state A.

The action leading to state A is preferred because its relevant state value estimate is 8, which is greater than 3.

leads toleads toDecisionAction to AState AV = 8Action to BState BV = 3
How do competing value estimates lead the agent to prefer one decision over another?

The same preference can be formed through action values, but the comparison is made in a different place. With V, the agent compares the values of the states reached by the available actions. With Q, the agent compares the estimates already attached to the actions themselves. In either case, the preferred choice is associated with the larger predicted total future reward.

What do you think happens?

Suppose two available actions have action-value estimates Q(action 1) = 6 and Q(action 2) = 4. Which action is preferred when choosing through action values?

  • Action 1
  • Action 2
  • Both are preferred because Q does not compare actions
  • There is not enough information to compare the estimates
Reveal answer

Answer: Action 1

When using action values, the agent favors the action with the highest action-value estimate. Here, 6 is greater than 4.

Mistakes with Value Notation

  • Treating q∗(a) and Qt(a) as the same quantity.

    q∗(a) denotes the true value, while Qt(a) denotes an estimate of that true value at time t.

    Fix: Use the star to recognize the true value and the subscript t to recognize the time-specific estimate.

  • Interpreting t as part of the action name.

    The symbol a identifies the action. The subscript t identifies when the estimate is being considered.

    Fix: Read Qt(a) as the estimate at time t for action a.

  • Confusing a state value with an action value.

    V assigns an estimate to a state, while Q assigns an estimate to a specific action.

    Fix: Ask what is being evaluated: the situation represented by a state, or a particular action.

  • Treating a value estimate as an immediate reward.

    The estimate is a prediction of total future reward available from the state, not a description limited to the next immediate result.

    Fix: Interpret the estimate as a long-run prediction.

Check Your Interpretation

MEDIUM

For each statement, identify whether it describes q∗(a), Qt(a), V, or Q. Then explain what future reward the quantity is predicting: the reward associated with a state generally, or the reward associated with a particular action.

Hints
  • Look for the star when identifying the true action value.
  • Look for the subscript t when identifying an estimate at a particular time.
  • V evaluates a state; Q evaluates a specific action.
  • Value estimates concern total future reward, not only the next immediate result.
  1. Identify whether the notation refers to a true value or an estimate.
  2. If it is an estimate, inspect the time index.
  3. Determine whether the quantity evaluates a state or a specific action.
  4. Compare the relevant estimates when deciding which choice is preferred.
  5. Interpret the larger estimate as the larger predicted total future reward.

Essential Distinctions

  1. q∗(a) denotes the true value of action a.
  2. Qt(a) denotes the estimate at time t of q∗(a); a identifies the action and t identifies the time of the estimate.
  3. A value estimate is a long-run prediction of total future reward.
  4. V evaluates a state generally, while Q evaluates a particular action.
  5. State-value choice compares the values of states reached by actions; action-value choice compares estimates attached directly to actions.

Key Takeaways

  • q∗(a) is the true value of action a, while Qt(a) is the estimate of that value at time t.
  • The time index t identifies when the estimate is being considered, not a different action.
  • Value estimates predict total future reward over the long run.
  • V describes the value of a state, whereas Q describes the value of taking a particular action.
  • An agent prefers the choice associated with the larger relevant value estimate.