Concepts / Rewards and Situations

Rewards and Situations

Reinforcement learning asks how to map situations to actions to maximize a numerical reward signal.

  • Programming

Learning Without Direct Instructions

Many learning problems begin with a direct instruction about what to do. Reinforcement learning begins differently: the learner must discover what to do by acting and observing the reward that follows. Its central question is how to map situations to actions so that a numerical reward signal is maximized.

The learner is not given a direct instruction identifying the best action. It tries actions, observes the resulting reward, and uses that information to learn.

The Situation-to-Action Mapping

A reinforcement learning problem involves connecting a situation with an action. When the learner encounters a situation, it selects an action. The outcome includes a numerical reward signal, which provides information about the consequence of that choice. Over time, the learner seeks a mapping from situations to actions that maximizes reward.

observedselectsproducesSituationcurrent inputLearnerselects an actionActionchosen responseRewardnumerical signal
How does the learner move from observing a situation to selecting an action and receiving a numerical reward?

A Simple Mapping

Suppose a learner encounters situation A and can try action X or action Y. How does reinforcement learning frame this choice?

Observe: The learner begins with situation A rather than with a direct instruction naming the correct action.

Choose: The learner selects one available action, such as action X or action Y.

Observe the outcome: The learner observes the numerical reward that follows the selected action.

Improve the mapping: The reward information helps the learner discover which action produces more reward in that situation.

The learner is building a situation-to-action mapping guided by numerical reward.

The Closed-Loop Interaction

Reinforcement learning problems are closed-loop because the learner's actions influence what happens next. An action does not merely produce a reward and end the process. It can influence the next situation, which becomes a later input for another decision.

selectsinfluencesproducesfeedsinformsLearnerchooses an actionActionselected responseNext situationlater inputRewardnumerical signalNext decisionanother action choice
How does a learner's action change the situation, produce a reward, and feed back into the next decision?

Closed-loop means that actions influence later inputs. The learner's choice changes the situation in which future choices are made.

Immediate and Later Consequences

An action can affect the reward received immediately and the rewards that become possible later. This happens because the action may influence the next situation, and that next situation can affect the rewards that follow. Therefore, evaluating an action may require considering its later consequences, not only its immediate result.

leads toaffectsinfluencesaffectsCurrent situationbefore actionActionchosen nowImmediate rewardreceived nowLater situationinfluenced by actionLater rewardsfollow future choices
How can one action affect both the reward received now and the rewards available in future situations?

Looking Beyond the Next Reward

Imagine two actions available in the same situation. One action produces a reward immediately but leads to a less rewarding later situation. The other produces a smaller immediate reward but leads to a situation with better later rewards. What must the learner consider?

Compare the immediate results: The learner observes the numerical reward that follows each action.

Trace the next situation: The learner considers how each action influences the next situation.

Consider later consequences: The learner considers how each resulting situation affects the rewards that follow.

The action with the larger immediate reward is not automatically the action that best supports maximizing reward over the interaction.

Discovering Through Trial and Error

Trial and error gives the learner a way to discover which actions produce the most reward. The learner tries an action, observes the reward that follows, and uses that experience when making later choices. Because the learner is not directly told the best action, repeated action and observation are central to discovering an effective situation-to-action mapping.

chooseproducesinformssupports another attemptSituationcurrent inputActionattemptRewardobserved resultLearningimproved choice information
How do repeated action attempts and their reward outcomes help the learner discover which actions work best?

Three Meanings of Reinforcement Learning

Reinforcement learning describes a problem: how to map situations to actions to maximize a numerical reward signal. It also describes a class of solution methods used for such problems, and the field that studies those problems and methods.

MeaningWhat it refers to
ProblemMapping situations to actions while maximizing a numerical reward signal.
Solution methodsMethods used to address those situation-to-action problems.
Field of studyThe broader study of the problems and the methods used to solve them.

The term reinforcement learning can refer to the problem, its solution methods, or the field studying both.

Mistakes About Rewards and Actions

  • Treating reinforcement learning as following a direct instruction.

    The learner must discover what to do by acting and observing the reward that follows.

    Fix: Think of the reward signal as information obtained after a choice, not as a direct instruction given before the choice.

  • Considering only the immediate reward.

    The action can influence the next situation, and that later situation can affect rewards that follow.

    Fix: Consider both the immediate result and the later consequences of the action.

  • Treating each decision as isolated.

    Reinforcement learning is closed-loop: actions influence later inputs.

    Fix: Trace how the action changes the next situation before evaluating the choice.

  • Using reward to mean only a verbal compliment or instruction.

    The central task is to maximize a numerical reward signal, and the learner uses the reward that follows action.

    Fix: Keep the numerical reward signal and the learner's discovered action choice conceptually separate.

Check Your Understanding

EASY

A learner takes an action and receives a numerical reward. The action also changes the next situation. Explain why the learner should not evaluate the action using only the reward received immediately.

Hints
  • Identify what the action changes besides the immediate reward.
  • Explain how the next situation can affect later rewards.
MEDIUM

In your own words, distinguish the three meanings of reinforcement learning: a problem, a class of solution methods, and a field of study.

Hints
  • Start with the situation-to-action mapping and numerical reward signal.
  • Then distinguish the methods for addressing the problem from the broader field that studies the problem and methods.

Key Takeaways

  1. Reinforcement learning maps situations to actions in order to maximize a numerical reward signal.
  2. The term refers to a problem, solution methods, and the field that studies both.
  3. The interaction is closed-loop because actions influence later inputs and situations.
  4. Trial and error lets a learner discover which actions produce the most reward.
  5. An action can affect both its immediate reward and the later rewards available through the situations it creates.

Key Takeaways

  • Reinforcement learning is concerned with mapping situations to actions to maximize a numerical reward signal.
  • The learner discovers useful actions through trial and error rather than receiving a direct instruction naming the best choice.
  • The problem is closed-loop because actions influence later situations and inputs.
  • Evaluating an action requires considering its immediate reward and its possible later consequences.
  • Reinforcement learning can mean the problem, the solution methods, or the field of study.