Immediate and Future Rewards
Discounting lets an agent compare rewards that arrive at different times.
The Time Problem
An agent may receive one reward now and another reward later. Comparing only the numerical sizes of those rewards is not enough when time matters. Discounting gives the agent a way to compare rewards that arrive at different times by assigning less present value to rewards that are farther away.
Following the Discount Rate
The discount rate γ controls how much present value remains in a future reward. Its allowed range is 0 ≤ γ ≤ 1. A lower γ makes behavior more focused on immediate rewards. As γ approaches 1, rewards arriving later have more influence on the agent's choice.
| Discount rate | Behavioral emphasis | Effect on later rewards |
|---|---|---|
| γ = 0 | Myopic | The agent focuses on the immediate reward |
| γ between 0 and 1 | Partly farsighted | The agent considers later rewards, but gives them less present influence |
| γ approaching 1 | Increasingly farsighted | Later consequences have more influence on the choice |
The discount rate determines how strongly the agent values rewards according to when they arrive.
A Choice With Consequences
Immediate Gain or Future Access
Imagine an agent choosing between two actions. Action A provides a larger reward immediately but reduces access to rewards later. Action B provides a smaller reward immediately and preserves access to a larger reward later.
Step 1: Consider only the next reward: An agent focused only on the immediate reward may choose Action A because its first reward is larger.
Step 2: Consider the later opportunity: An agent with a discount rate closer to 1 gives more influence to the later reward that Action B keeps available.
Step 3: Compare discounted returns: The agent should compare the rewards after discounting the later outcome, rather than choosing solely from the immediate reward.
The action with the largest immediate reward is not automatically the action with the largest discounted return.
Comparing Now and Later
To compare a reward received now with one received several steps later, first adjust the later reward to its present value through discounting. The farther away the reward is, the less present value discounting assigns to it. The discount rate determines how strong that reduction is.
| Question | Immediate-reward view | Discounted-return view |
|---|---|---|
| What is considered first? | The reward available now | Rewards now and later |
| What happens to a delayed reward? | It may be ignored | Its present value is reduced according to γ and its delay |
| What kind of choice can result? | A choice favoring the largest immediate reward | A choice that accounts for future consequences |
Myopic and Farsighted Choices
At γ = 0, behavior is myopic: the agent focuses on the next reward. This can favor an action that looks best immediately even when that action reduces future opportunities. As γ approaches 1, behavior becomes increasingly farsighted because later consequences receive more influence in the decision.
Assuming the largest immediate reward must produce the best overall result.
The immediate reward is only one part of the discounted return.
Fix:
Consider the later rewards that each action makes available, then compare their discounted present values.Treating γ as the size of the reward.
γ controls how much present value remains in a future reward; it does not describe the reward itself.
Fix:
Use γ to describe the agent's emphasis on immediate versus delayed rewards.Thinking that a delayed reward has the same present value regardless of when it arrives.
Discounting assigns less present value to rewards that are farther away.
Fix:
Account for both the reward's timing and the chosen discount rate.
Practice the Comparison
An agent can choose Action A, which gives a larger reward immediately but reduces access to later rewards, or Action B, which gives a smaller immediate reward and preserves later access. Explain which action a myopic agent is more likely to prefer and why an agent with γ closer to 1 may evaluate the actions differently.
Hints
- Start with the meaning of γ = 0.
- Identify what happens to later opportunities after Action A.
- Then consider why increasing γ gives later consequences more influence.
- A strong answer should say that the myopic agent is more likely to choose Action A because it emphasizes the immediate reward. An agent with γ closer to 1 gives more weight to the later rewards preserved by Action B, so Action B may produce the larger discounted return even though its first reward is smaller.
Key Takeaways
- Discounting lets an agent compare rewards that arrive at different times.
- The discount rate γ controls how much present value remains in a future reward, with 0 ≤ γ ≤ 1.
- γ = 0 produces myopic behavior focused on the immediate reward.
- As γ approaches 1, the agent becomes increasingly farsighted and gives later consequences more influence.
- Choosing the largest immediate reward can reduce future opportunities and therefore reduce total discounted return.