Concepts / Rewards and Penalties in Reinforcement Learning

Rewards and Penalties in Reinforcement Learning

Exploration tests unfamiliar actions so the agent can improve future decisions.

  • Programming

The Choice Behind Every Action

A reinforcement learning agent faces a practical tension whenever it chooses an action. It can repeat an action that has already produced reward, or it can try an action it has not selected before in case that action turns out to be better. These two choices are called exploitation and exploration.

testsusesExplorationunfamiliar actionsNew evidencepossible future improvementExploitationeffective actionsObtained rewarduse current evidence
How does an agent balance trying unfamiliar actions with repeating actions that have worked before?

Tracing a Decision

Imagine an agent choosing between two actions. One action has already produced reward, while the other has not yet been selected. Choosing the known action is exploitation: the agent uses evidence already collected to obtain reward. Choosing the unfamiliar action is exploration: the agent gathers new evidence that may improve later decisions.

try another optionproducesinformsKnown actioncurrent evidenceUnfamiliar actionexplorationReward or penaltynew evidenceLater actionupdated preference
What happens next when an agent explores a new action, receives feedback, and updates which action it chooses?

The important sequence is not that exploration permanently replaces exploitation. Exploration supplies new evidence about possible actions. Exploitation uses the evidence already collected to obtain reward. The agent needs both behaviors because they advance different parts of the task.

Balancing Evidence and Reward

Exploration helps an agent discover actions that may deserve preference later. Exploitation helps the agent benefit from actions that already appear effective. A useful balance therefore contains enough exploration to discover effective actions and enough exploitation to use what has been discovered.

Choice patternWhat it usesPotential failure
Only exploitationCurrent experienceThe agent does not try actions that could lead to better future choices.
Only explorationUnfamiliar actionsThe agent does not make use of actions that past experience has shown to be effective.
Both behaviorsNew and existing evidenceThe agent can discover possible improvements while benefiting from known effective actions.

When Outcomes Vary

The dilemma becomes especially demanding when outcomes are stochastic, meaning that outcomes vary. In that situation, one experience may not be enough to judge an action confidently. The agent may need to try an action many times before obtaining a reliable estimate of its expected reward.

Choosing Between Two Uncertain Actions

An agent has one action with past evidence of reward and another action that has not been selected before. The result of either action can vary from trial to trial.

Start with current evidence: The agent can exploit the action that already appears effective, but that evidence may not reveal whether the unfamiliar action could be better.

Explore the unfamiliar action: The agent tests the unfamiliar action to collect new evidence rather than judging it without experience.

Repeat when necessary: Because outcomes vary, repeated trials may be needed before the agent can judge the action reliably.

Use the growing evidence: The agent can progressively favor actions that appear best while still recognizing that discovery required exploration.

A single result should not automatically settle the comparison when outcomes vary; repeated trials can provide a more reliable basis for later choices.

Reward at Time t

R_t denotes the reward signal at time t. It is a theoretical abstraction used with the environment to help define the reinforcement learning problem an agent is trying to solve.

next time stepnext time stepTime 1R₁Time 2R₂Time 3R₃
How does the reward signal R_t relate to successive time steps as an agent acts and receives feedback?

The subscript t identifies the time at which the signal is considered. The notation does not by itself claim that the signal is a physical object in the agent, nor does it say that every biological learning system contains one unchanged signal with this exact form.

A reward signal helps define the problem together with the environment. It is part of the model-level description of what the agent is trying to solve.

Reward and Reinforcement

Reward signals and reinforcement signals are distinct concepts. In reinforcement learning theory, R_t is the theoretical reward signal at time t. Reinforcement signals should not simply be treated as another name for that theoretical symbol.

helps definerelates toReward signaltheoretical R_tProblem definitionmodel levelReinforcementsignaldistinct conceptBehavior changelearning process
What is the difference between the theoretical reward signal used to define an RL problem and the broader reinforcement concept?

From Models to Brains

Neural signals are physiological events that may behave like theoretical signals in function. This means a theoretical model can describe a useful role with one symbol even when the biological system involves many neural signals and many systems.

usesmay involvecontribute todescribes a functionRL modelR_tDefined problemtheoretical descriptionNeural signalsphysiological eventsMany systemsbiological organizationLearningfunctional relationship
How can multiple physiological signals contribute to learning without corresponding to one literal, unitary master reward signal R_t?

When interpreting R_t in a biological context, describe it first as a theoretical reward signal. Do not automatically identify it with one specific physiological event in an animal's brain.

  • Treating R_t as a literal unitary master reward signal inside an animal's brain.

    R_t is a theoretical abstraction, while biological systems may involve many neural signals and many systems.

    Fix: Use R_t as the model-level reward signal at time t, and separately discuss physiological signals as events that may behave similarly in function.

  • Using exploration and exploitation as if they were independent strategies.

    Exploration supplies evidence and exploitation uses evidence; the agent needs both behaviors.

    Fix: Explain how the two behaviors contribute different parts of the same decision problem.

  • Assuming that one trial reliably identifies the best action when outcomes vary.

    A single experience may not be enough to estimate an action's expected reward reliably.

    Fix: Recognize that repeated trials may be needed before progressively favoring actions that appear best.

Apply the Dilemma

MEDIUM

An agent has repeatedly received reward after choosing one action. A second action has not yet been selected. Explain what the agent gains from exploiting the first action, what it gains from exploring the second action, and why using only one of these behaviors can be insufficient.

Hints
  • Connect exploitation with evidence already collected.
  • Connect exploration with new evidence about a possible better action.
  • Consider what each extreme would prevent the agent from doing.

What do you think happens?

If outcomes vary from trial to trial, is one experience enough to judge an unfamiliar action confidently?

  • Yes, one experience is always enough.
  • No, repeated trials may be needed.
  • Only exploitation can provide evidence.
Reveal answer

Answer: No, repeated trials may be needed.

On a stochastic task, repeated trials can provide a more reliable estimate of an action's expected reward.

Key Takeaways

  1. Exploration tests unfamiliar actions and supplies new evidence for future decisions.
  2. Exploitation uses actions already known to be effective to obtain reward.
  3. An agent needs both behaviors because only exploration ignores useful past evidence, while only exploitation prevents discovery of potentially better actions.
  4. R_t is the theoretical reward signal at time t and helps define the reinforcement learning problem together with the environment.
  5. R_t should not be treated as one literal, unchanged physiological reward signal in an animal's brain.

Key Takeaways

  • Exploration tries unfamiliar actions; exploitation repeats actions that past experience suggests are effective.
  • The two behaviors are complementary: exploration discovers possibilities, while exploitation uses current knowledge to obtain reward.
  • Exclusive exploration and exclusive exploitation can each cause failure in different ways.
  • R_t denotes the theoretical reward signal at time t and helps define the agent's problem.
  • A model-level reward signal should not be identified automatically with one unitary physiological signal in an animal's brain.