Concepts / Models of Classical Conditioning

Models of Classical Conditioning

The Rescorla-Wagner model links learning to surprise.

  • Programming

Why Conditioning Needs a Model

Classical conditioning is not only a matter of recording which events occur together. A useful model must also explain how learning changes from one trial to the next. Prior conditioning can alter what an animal learns later, and the time at which information becomes available can affect how predictions change during a conditioning episode.

The central comparison in this article is between the Rescorla-Wagner emphasis on surprise and the TD model's emphasis on temporal changes in prediction.

Reinforcement in Historical Context

The word reinforcement has not always had one fixed meaning in animal learning. According to the source, the term first appeared, to the best of its knowledge, in the 1927 English translation of Pavlov's monograph. In that use, reinforcement referred to a non-action-contingent case: the event was not described as depending on an action.

Later traditions used the term differently. Mackintosh proposed using reinforcement for either strengthening or weakening a pattern of behavior. Skinner used reinforcement only for strengthening behavior and treated weakening as something produced by punishment. These differences are historically important because the same word can point to different ideas depending on the theoretical tradition.

Historical usageMeaning described in the source
Pavlov translationA non-action-contingent event in the context of animal learning
MackintoshEither strengthening or weakening a pattern of behavior
SkinnerStrengthening behavior; weakening was treated as produced by punishment

When reading a conditioning theory, check what the author means by reinforcement before comparing that theory with another one. A shared term does not guarantee a shared concept.

Kamin Blocking Across Two Stages

Kamin blocking shows that prior conditioning can affect subsequent learning. The situation has two stages. First, one stimulus is repeatedly involved in conditioning. Later, that previously conditioned stimulus appears together with a new stimulus. The question is whether the new stimulus will also become a predictor of the outcome. In blocking, prior conditioning prevents or greatly reduces the expected new learning about the added stimulus.

prior conditioningpredictsappears alongsidealready expected, so Cue B gains littleCue AconditioningCue Aalready predicts outcomeOutcomeexpectedCue Blittle new learningCue Bintroduced with Cue A
What changes when a new cue is introduced alongside a cue that already predicts the outcome, and why is the new cue learned weakly?

Tracing a blocking scenario

A learner first experiences Cue A repeatedly in a conditioning arrangement. Later, Cue A and a new Cue B appear together. Explain why Cue B may show little new learning.

First stage: Cue A is repeatedly involved in conditioning, so the learner develops a prior conditioning history for that stimulus.

Second stage: Cue A appears together with the new Cue B. The important comparison is not just that both cues are present, but what the learner already expects because of Cue A.

Expected outcome: Because Cue A already predicts the outcome, the later presentation supplies less unexpected information than it would have supplied without the earlier conditioning.

Result for Cue B: Kamin blocking describes the case in which the new Cue B is learned weakly or not as expected because prior conditioning has changed the conditions for later learning.

Prior learning about Cue A changes the outcome of later learning about Cue B.

What do you think happens?

After Cue A has already been conditioned, what would the blocking account predict about learning for a new Cue B presented together with Cue A?

  • Cue B should necessarily receive the same learning as Cue A
  • Cue B may receive little or no additional learning
  • Prior conditioning should have no effect on Cue B
Reveal answer

Answer: Cue B may receive little or no additional learning.

Kamin blocking is defined in the source as a case in which prior conditioning affects subsequent learning. The previously conditioned cue makes the outcome less surprising when the compound is presented.

Surprise in Rescorla-Wagner

The Rescorla-Wagner model links learning to surprise. Learning occurs when animals are surprised, so the amount of new learning depends on the difference between what the learner expects and what occurs.

This idea gives blocking a learning-based explanation. Before the compound stage, Cue A has already acquired predictive significance. When Cue A and Cue B appear together, the outcome is not as unexpected as it would have been for a learner with no prior conditioning. Since the outcome is already prepared for, the new Cue B may contribute little or no additional learning.

unexpected outcomechanges expectationless unexpected informationprior conditioning reduces surpriseEarly trialoutcome less expectedMore new learninggreater surpriseLater trialexpectation has changedLess new learningless surpriseCompound trialprior cue already predictsoutcomeWeak Cue B learninglittle additional learning
How does the difference between the expected and actual outcome change the amount learned on each conditioning trial?

To analyze a Rescorla-Wagner explanation, ask two questions: What does the learner already predict, and how much unexpected information remains when the outcome occurs?

Alternative Conditioning Models

The Rescorla-Wagner model is one model among several used to study classical conditioning. The source lists models associated with Klopf, Grossberg, Mackintosh, Moore and Stickney, Pearce and Hall, and Courville, Daw, and Touretzky. This broader list matters because it places Rescorla and Wagner within a history of animal learning theory rather than treating their model as the only available explanation.

Model or model family named in the sourceWhat can be stated from the source pack
KlopfAn alternative model associated with the study of classical conditioning
GrossbergAn alternative model associated with the study of classical conditioning
MackintoshA model associated with the broader history of animal learning theory
Moore and StickneyAn alternative model associated with the study of classical conditioning
Pearce and HallAn alternative model associated with the study of classical conditioning
Courville, Daw, and TouretzkyAn alternative model associated with the study of classical conditioning

The source identifies these alternatives but does not specify their mechanisms in the supplied material.

Temporal Differences in a Conditioning Episode

The TD model places time and changes in prediction at the center of its account of classical conditioning. A temporal difference concerns the relationship between predictions at different moments in a conditioning episode.

Imagine following one conditioning episode from its beginning to its end. At the beginning, the learner has one set of expectations. As the episode unfolds, later events provide new information. The TD perspective compares these successive points rather than treating the whole episode as one undivided event. Learning is therefore connected to changes in predictive information over time.

time advancesnew predictive informationcompare neighboring momentsEpisode beginninginitial predictionCue appearsnew informationUpdated predictionprediction changesOutcomelater information
How does the prediction error change from one moment or trial to the next as the expected outcome becomes more accurate?

Reading one episode as temporal change

Explain a simple conditioning episode from a TD perspective without treating the entire episode as one undifferentiated event.

Start with the beginning: Describe the learner's expectations at the beginning of the episode.

Advance to the cue: When the cue appears, it provides information at a particular moment. The model considers how this changes the predictive situation.

Advance to later information: As the episode continues toward the outcome, the learner receives further information and can compare the prediction at this moment with the prediction at an earlier moment.

Interpret the difference: The learning-relevant quantity is the time-sensitive change in predictive information, not merely the fact that the cue and outcome occurred in the same episode.

The TD account organizes conditioning around successive predictions and the differences between them.

A temporal difference is not simply a label for the final outcome. It identifies a time-sensitive change in predictive information that the model uses to organize learning.

Three Learning Rules in One History

The TD model developed within a larger history of learning rules. Sutton and Barto's 1981 work recognized a near identity between the Rescorla-Wagner model and the Least-Mean-Square learning rule, also called the LMS or Widrow-Hoff rule. Widrow and Hoff introduced the LMS rule in 1960. This connection places the TD model alongside an established error-based learning-rule tradition rather than presenting it as an unrelated account of conditioning.

learning-rule traditionnear identity recognized with LMSshared historical contextLMS ruleWidrow-Hoff, 1960Rescorla-Wagner modelconditioning modelTD modelSutton and Barto historyNear identityrecognized in 1981
How are the TD model, Rescorla-Wagner model, and LMS learning rule connected historically and mathematically?
ItemHistorical point
LMS or Widrow-Hoff ruleIntroduced by Widrow and Hoff in 1960
Rescorla-Wagner and LMSSutton and Barto's 1981 work recognized a near identity between them
Early TD modelAn early version appeared in Sutton and Barto in 1981
TD algorithmSutton later developed the algorithm in 1984 and 1988
TD model presentationsThe model was revised and presented in Sutton and Barto in 1987, with a more complete presentation in 1990

The source describes a gradual development rather than one finished theory appearing in a single publication.

The development was gradual. The early version of the TD model made a specific prediction: temporal primacy overrides blocking. The source reports that Kehoe, Scheurs, and Graham later showed this effect in the rabbit nictitating membrane preparation. This is an example of a model being evaluated against an experimental result, not merely a restatement that timing exists.

Why Representation Changes the Evaluation

The TD model does not depend only on its learning rule. The way a stimulus is represented can also affect how the model performs. Later work examined different stimulus representations and their possible neural implementations in relation to response timing and response topography.

A microstimulus representation represents a stimulus through time-specific microstimuli across an episode. The source emphasizes that this representation should not be assigned one fixed physical form automatically; it is a way of examining how stimulus information is encoded across time.

represented over timerepresented over timerepresented over timesupports a time-specific predictionsupports a later predictionSingle stimulusone episodeMicrostimulus 1early momentPrediction 1based on earlyrepresentationMicrostimulus 2middle momentPrediction 2based on laterrepresentationMicrostimulus 3later moment
How does representing a single stimulus as a sequence of time-specific microstimuli change which predictions and errors the model computes?

A temporal model needs a way to encode stimulus information across the episode. If the representation preserves useful distinctions between early and later moments, the model has a richer basis for comparing successive predictions. If those distinctions are not represented, the model has less information with which to organize time-sensitive changes.

Ludvig, Sutton, and Kehoe introduced the microstimulus representation in 2008. In 2012, they evaluated the TD model on previously unexplored classical-conditioning tasks and examined the influence of several stimulus representations, including the microstimulus representation. The important lesson is that evaluating a temporal learning rule also requires asking how the stimulus is represented.

Same stimulus, different representational question

Why might a TD evaluation ask whether a stimulus is represented as one undifferentiated event or through time-specific microstimuli?

Identify the TD requirement: The TD model compares predictive information at different moments, so it needs stimulus information that can be related to those moments.

Consider one undivided representation: If the stimulus is treated without useful time-specific distinctions, the model has a less detailed basis for comparing early and later predictions.

Consider microstimuli: A microstimulus representation examines the stimulus through a sequence of time-specific components, allowing the evaluation to ask which components support which predictions.

Draw the evaluation lesson: Differences in performance may reflect not only the learning rule but also the representation supplied to that rule.

Stimulus representation is part of the explanation being tested, not merely an implementation detail outside the model.

Common Reasoning Errors

  • Treating conditioning trials as independent events

    Kamin blocking specifically depends on prior conditioning changing the conditions for later learning.

    Fix: Trace the learning history before interpreting what happens when the new stimulus is introduced.

  • Defining learning as simple stimulus-outcome co-occurrence

    The Rescorla-Wagner account emphasizes whether the outcome is surprising given what is already predicted.

    Fix: Ask what the learner already expects before deciding how much new learning remains.

  • Treating the TD model as a model of only the final outcome

    The TD model focuses on changes in predictive information across different moments.

    Fix: Compare neighboring moments and ask how the prediction changes as new information becomes available.

  • Assuming the TD learning rule is sufficient without specifying a representation

    The source states that stimulus representation, including microstimulus representation, affects how the TD model performs.

    Fix: Treat representation as part of the evaluation of the model.

  • Assuming reinforcement has one universal historical meaning

    The source describes meaningful differences among these traditions.

    Fix: Identify the historical tradition and its definition before comparing claims.

  • Inventing mechanisms for every alternative model named in the source

    The source lists these alternatives but does not provide their detailed mechanisms here.

    Fix: Use the supplied material only to identify them as part of the broader history of conditioning models.

Practice the Model Trace

MEDIUM

A learner first experiences Cue A in a conditioning arrangement. Later, Cue A is presented together with a new Cue B. Write a short explanation of the result using both of these perspectives: first, explain why Cue B may receive little learning under the Rescorla-Wagner emphasis on surprise; second, explain what a TD analysis would examine across the episode.

Hints
  • Begin with the learner's prior conditioning history for Cue A.
  • For the Rescorla-Wagner perspective, ask whether the outcome is already expected.
  • For the TD perspective, identify successive moments and the changes in predictive information between them.
  • Mention why the representation of the stimulus may matter for the TD analysis.

A complete comparison

Compare the Rescorla-Wagner and TD explanations of the same two-stage conditioning situation.

State the prior learning: Cue A was conditioned in the first stage, so the learner enters the second stage with an established expectation associated with Cue A.

Apply the Rescorla-Wagner emphasis: When Cue A and Cue B appear together, the outcome is less surprising because Cue A already provides predictive information. Therefore Cue B may receive little or no additional learning.

Apply the TD emphasis: Follow the episode across time and compare predictions at successive moments instead of treating the compound and outcome as one undivided event.

Add representation: Ask how the stimulus is encoded across the episode. A time-sensitive representation, such as the microstimulus representation examined in later work, gives the model a basis for evaluating predictions at different moments.

The Rescorla-Wagner account emphasizes surprise and prior expectation, while the TD account emphasizes time-dependent changes in prediction and the representation that makes those changes available.

Key Takeaways

  1. The meaning of reinforcement changed across historical traditions in animal learning, so the term must be interpreted in context.
  2. Kamin blocking demonstrates that prior conditioning can affect later learning about a newly added stimulus.
  3. The Rescorla-Wagner model links learning with surprise: when an outcome is already expected, additional learning may be small.
  4. The TD model explains conditioning through changes in predictive information across time and developed within a history connected to the Rescorla-Wagner model and the LMS learning rule.
  5. Stimulus representation matters for TD evaluations, and microstimulus representation provides a way to examine stimulus information through time-specific components.

Key Takeaways

  • Prior conditioning changes the conditions for later learning, as shown by Kamin blocking.
  • The Rescorla-Wagner model emphasizes surprise and explains why an already expected outcome may produce little new learning.
  • The TD model compares predictions across moments in a conditioning episode rather than treating the episode as a single undivided event.
  • The TD model has a historical relationship with the Rescorla-Wagner model and the LMS or Widrow-Hoff learning rule.
  • Stimulus representation, including microstimulus representation, affects how the TD model can be evaluated on conditioning tasks.