Concepts / Immediate and Future Rewards

Immediate and Future Rewards

Discounting lets an agent compare rewards that arrive at different times.

  • Programming

The Time Problem

An agent may receive one reward now and another reward later. Comparing only the numerical sizes of those rewards is not enough when time matters. Discounting gives the agent a way to compare rewards that arrive at different times by assigning less present value to rewards that are farther away.

more time passesmore time passesReward nowhighest present valueReward laterless present valueReward farther awaystill less present value
How does the value assigned to the same reward change as the number of time steps before receiving it increases?

Following the Discount Rate

The discount rate γ controls how much present value remains in a future reward. Its allowed range is 0 ≤ γ ≤ 1. A lower γ makes behavior more focused on immediate rewards. As γ approaches 1, rewards arriving later have more influence on the agent's choice.

Discount rateBehavioral emphasisEffect on later rewards
γ = 0MyopicThe agent focuses on the immediate reward
γ between 0 and 1Partly farsightedThe agent considers later rewards, but gives them less present influence
γ approaching 1Increasingly farsightedLater consequences have more influence on the choice

The discount rate determines how strongly the agent values rewards according to when they arrive.

increase γincrease γγ = 0immediate reward dominates0 < γ < 1later reward has someinfluenceγ approaching 1later consequences mattermore
How do different values of γ, from 0 to 1, change the relative value of immediate versus delayed rewards?

A Choice With Consequences

Immediate Gain or Future Access

Imagine an agent choosing between two actions. Action A provides a larger reward immediately but reduces access to rewards later. Action B provides a smaller reward immediately and preserves access to a larger reward later.

Step 1: Consider only the next reward: An agent focused only on the immediate reward may choose Action A because its first reward is larger.

Step 2: Consider the later opportunity: An agent with a discount rate closer to 1 gives more influence to the later reward that Action B keeps available.

Step 3: Compare discounted returns: The agent should compare the rewards after discounting the later outcome, rather than choosing solely from the immediate reward.

The action with the largest immediate reward is not automatically the action with the largest discounted return.

choose Action Areduceschoose Action BpreservesAction choiceImmediate gainlarger nowLater accessreducedImmediate gainsmaller nowLater accesspreserved
What happens to the sequence of available rewards when an agent takes a smaller immediate reward instead of choosing an action that reduces later access?

Comparing Now and Later

To compare a reward received now with one received several steps later, first adjust the later reward to its present value through discounting. The farther away the reward is, the less present value discounting assigns to it. The discount rate determines how strong that reduction is.

express at presentdiscount for delaycompare present valuesReward noworiginal timingPresent value nowimmediate rewardReward lateroriginal timingPresent value laterdiscounted reward
How can an agent compare a reward received now with a reward received several steps later after adjusting both to present value?
QuestionImmediate-reward viewDiscounted-return view
What is considered first?The reward available nowRewards now and later
What happens to a delayed reward?It may be ignoredIts present value is reduced according to γ and its delay
What kind of choice can result?A choice favoring the largest immediate rewardA choice that accounts for future consequences

Myopic and Farsighted Choices

prioritizesconsiders more stronglyγ = 0choose by immediate rewardImmediate outcomestrong influence for bothγ approaching 1give later consequencesmore influenceLater outcomemore influence as γ rises
How does the agent’s choice change when it values only the next reward compared with valuing rewards far into the future?

At γ = 0, behavior is myopic: the agent focuses on the next reward. This can favor an action that looks best immediately even when that action reduces future opportunities. As γ approaches 1, behavior becomes increasingly farsighted because later consequences receive more influence in the decision.

  • Assuming the largest immediate reward must produce the best overall result.

    The immediate reward is only one part of the discounted return.

    Fix: Consider the later rewards that each action makes available, then compare their discounted present values.

  • Treating γ as the size of the reward.

    γ controls how much present value remains in a future reward; it does not describe the reward itself.

    Fix: Use γ to describe the agent's emphasis on immediate versus delayed rewards.

  • Thinking that a delayed reward has the same present value regardless of when it arrives.

    Discounting assigns less present value to rewards that are farther away.

    Fix: Account for both the reward's timing and the chosen discount rate.

Practice the Comparison

MEDIUM

An agent can choose Action A, which gives a larger reward immediately but reduces access to later rewards, or Action B, which gives a smaller immediate reward and preserves later access. Explain which action a myopic agent is more likely to prefer and why an agent with γ closer to 1 may evaluate the actions differently.

Hints
  • Start with the meaning of γ = 0.
  • Identify what happens to later opportunities after Action A.
  • Then consider why increasing γ gives later consequences more influence.
  1. A strong answer should say that the myopic agent is more likely to choose Action A because it emphasizes the immediate reward. An agent with γ closer to 1 gives more weight to the later rewards preserved by Action B, so Action B may produce the larger discounted return even though its first reward is smaller.

Key Takeaways

  • Discounting lets an agent compare rewards that arrive at different times.
  • The discount rate γ controls how much present value remains in a future reward, with 0 ≤ γ ≤ 1.
  • γ = 0 produces myopic behavior focused on the immediate reward.
  • As γ approaches 1, the agent becomes increasingly farsighted and gives later consequences more influence.
  • Choosing the largest immediate reward can reduce future opportunities and therefore reduce total discounted return.