Concepts / Reward Prediction and Animal Learning

Reward Prediction and Animal Learning

Classical conditioning is defined by the independence of the reinforcing stimulus from the animal's behavior.

  • Programming

Prediction or Control

Animal learning can involve two different problems. In a control problem, the learner must act in a way that produces a reinforcing stimulus. In a prediction problem, the learner learns that one event signals another. The distinction between instrumental conditioning and classical conditioning follows from this difference: instrumental conditioning concerns behavior that produces reinforcement, whereas classical conditioning concerns predictive information about events.

requiresproduceslearns aboutpredictsInstrumentalconditioningAnimal actionReinforcing stimulusClassicalconditioningPredictive eventReinforcing stimulus
Does the animal's behavior determine whether the reinforcing stimulus occurs?

The diagnostic question is whether the reinforcing stimulus depends on the animal's behavior. If behavior produces the stimulus, the learning problem is about control. If the stimulus occurs independently of behavior and an event comes to predict it, the learning problem is the classical-conditioning pattern.

Tracing Event A

A predictor without an action

Suppose event A is followed by a reinforcing stimulus, even though the animal's behavior does not determine whether that stimulus appears.

Identify the dependency: The reinforcing stimulus is independent of the animal's behavior, so the learner is not being asked to discover an action that produces it.

Identify the learning target: The learner must identify event A as a predictor of the later reinforcing stimulus.

Classify the pattern: Because the central problem is learning that one event signals another, this is the basic classical-conditioning connection.

The learner acquires predictive information about event A rather than control over whether the reinforcing stimulus occurs.

signalsabout later eventEvent APredictionReinforcing stimulus
How does a learner acquire predictive information when the reinforcing stimulus does not depend on an action?

TD Prediction Across Time

Temporal-difference learning provides a reinforcement-learning account of prediction that parallels classical conditioning. In this use of TD learning, the learner is not primarily selecting an action to produce reinforcement. Instead, it is learning predictive information about events. The important addition is temporal structure: the model considers events within an individual trial, so the learner's prediction can be related to the position of an event within that trial.

next momentlater eventMoment 1Event AMoment 2PredictionMoment 3Reinforcing stimulus
How does predicted reward change as a cue and reinforcing stimulus unfold from one moment or state to the next?

The timeline matters because TD learning treats a trial as containing successive moments or states rather than as one undivided event. A cue can therefore be considered in relation to what follows it within the same trial. This is why learning to predict with TD learning corresponds to the classical-conditioning pattern: the model tracks predictive relations between events, while its temporal structure represents where those events occur in the unfolding trial.

What do you think happens?

If the reinforcing stimulus does not depend on the animal's action, what kind of information is the TD learner acquiring in this setting?

  • A way to produce the reinforcing stimulus through an action
  • A prediction that one event signals another
  • A rule for ignoring events within a trial
Reveal answer

Answer: A prediction that one event signals another

The classical-conditioning parallel is prediction. The TD model represents predictive relations between events and adds information about their timing within a trial.

Rescorla-Wagner and TD

ModelRepresentation emphasizedTreatment of a trial
Rescorla-Wagner accountLearning about eventsDoes not represent the trial's internal temporal dimension in the TD sense
TD modelLearning about events with temporal structureConsiders events within individual trials and their position in time

The TD model is described as a generalization of the influential Rescorla-Wagner model. The key extension is not a change from learning to prediction into learning to control. Both are relevant to learning about events. Instead, TD adds temporal structure by considering events within individual trials. This lets the model represent the sequence and timing of intermediate events rather than treating a trial as lacking internal timing.

representsrepresentsRescorla-WagnerTrialevents without internaltimingTD modelTrialevents across moments
What information about a trial does the TD model add to the Rescorla-Wagner account?

Second-Order Conditioning

Second-order conditioning describes how a predictor of a reinforcing stimulus can itself become reinforcing. The important sequence is that a new cue precedes an already learned predictor. The new cue does not need to be directly paired with the reinforcing stimulus in order to acquire significance; its relationship to the established predictor can give it predictive value.

precedespredictsNew cueLearned predictorReinforcing stimulus
How can a new cue acquire predictive value by preceding an already learned cue without being directly paired with the reinforcing stimulus?

A cue before event A

Event A already predicts a reinforcing stimulus. A new cue is placed before event A, but the new cue is never directly paired with the reinforcing stimulus.

Start with the established relation: Event A is already a predictor of the reinforcing stimulus.

Add the new cue: The new cue precedes event A, placing it earlier in the predictive sequence.

Follow the chain: The new cue is connected to event A, and event A is connected to the reinforcing stimulus.

Identify the second-order effect: The new cue can acquire predictive value and can itself become reinforcing, even without a direct pairing with the reinforcing stimulus.

A predictor of a predictor can become significant through its position in the predictive chain.

Common Misunderstandings

  • Treating classical conditioning as behavior that produces reinforcement.

    Classical conditioning is defined by the independence of the reinforcing stimulus from the animal's behavior.

    Fix: Ask whether the learner is acquiring control through action or prediction about events.

  • Describing TD learning in this context as primarily an action-selection account.

    The relevant reinforcement-learning parallel is TD learning used for prediction.

    Fix: Focus on how the learner predicts a later reinforcing stimulus from an earlier event.

  • Treating the TD model and the Rescorla-Wagner model as identical.

    The TD model extends the Rescorla-Wagner account by adding temporal structure.

    Fix: Represent the events and intermediate moments within the trial.

  • Assuming second-order conditioning requires direct pairing of every cue with the reinforcing stimulus.

    Second-order conditioning concerns a new cue that precedes an already learned predictor.

    Fix: Trace the predictive chain from the new cue to the learned predictor and then to the reinforcing stimulus.

Check Your Understanding

MEDIUM

A reinforcing stimulus appears whether or not the animal acts. Before it appears, event A reliably occurs. Later, a new cue is placed before event A. Explain which parts of this situation illustrate classical conditioning, TD prediction, temporal structure, and second-order conditioning.

Hints
  • First determine whether the reinforcing stimulus depends on behavior.
  • Then identify what event the learner is predicting.
  • Ask what the TD model adds when the events are considered within one trial.
  • Finally, trace how the new cue relates to event A and the reinforcing stimulus.
  1. Classical conditioning is a prediction problem in which the reinforcing stimulus is independent of the animal's behavior. TD learning parallels this pattern when it is used to learn predictions about events rather than to select actions that produce reinforcement. The TD model extends the Rescorla-Wagner account by representing temporal structure within individual trials. Second-order conditioning follows when a new cue precedes an already learned predictor and thereby acquires predictive value itself.

Key Takeaways

  • Classical conditioning concerns prediction when reinforcement is independent of behavior.
  • Instrumental conditioning concerns control because behavior produces the reinforcing stimulus.
  • TD learning provides a reinforcement-learning account of prediction that parallels classical conditioning.
  • TD extends the Rescorla-Wagner model by representing events and their timing within individual trials.
  • Second-order conditioning allows a new cue to gain predictive value by preceding an already learned predictor.