Models of Classical Conditioning
The Rescorla-Wagner model links learning to surprise.
Why Conditioning Needs a Model
Classical conditioning is not only a matter of recording which events occur together. A useful model must also explain how learning changes from one trial to the next. Prior conditioning can alter what an animal learns later, and the time at which information becomes available can affect how predictions change during a conditioning episode.
The central comparison in this article is between the Rescorla-Wagner emphasis on surprise and the TD model's emphasis on temporal changes in prediction.
Reinforcement in Historical Context
The word reinforcement has not always had one fixed meaning in animal learning. According to the source, the term first appeared, to the best of its knowledge, in the 1927 English translation of Pavlov's monograph. In that use, reinforcement referred to a non-action-contingent case: the event was not described as depending on an action.
Later traditions used the term differently. Mackintosh proposed using reinforcement for either strengthening or weakening a pattern of behavior. Skinner used reinforcement only for strengthening behavior and treated weakening as something produced by punishment. These differences are historically important because the same word can point to different ideas depending on the theoretical tradition.
| Historical usage | Meaning described in the source |
|---|---|
| Pavlov translation | A non-action-contingent event in the context of animal learning |
| Mackintosh | Either strengthening or weakening a pattern of behavior |
| Skinner | Strengthening behavior; weakening was treated as produced by punishment |
When reading a conditioning theory, check what the author means by reinforcement before comparing that theory with another one. A shared term does not guarantee a shared concept.
Kamin Blocking Across Two Stages
Kamin blocking shows that prior conditioning can affect subsequent learning. The situation has two stages. First, one stimulus is repeatedly involved in conditioning. Later, that previously conditioned stimulus appears together with a new stimulus. The question is whether the new stimulus will also become a predictor of the outcome. In blocking, prior conditioning prevents or greatly reduces the expected new learning about the added stimulus.
Tracing a blocking scenario
A learner first experiences Cue A repeatedly in a conditioning arrangement. Later, Cue A and a new Cue B appear together. Explain why Cue B may show little new learning.
First stage: Cue A is repeatedly involved in conditioning, so the learner develops a prior conditioning history for that stimulus.
Second stage: Cue A appears together with the new Cue B. The important comparison is not just that both cues are present, but what the learner already expects because of Cue A.
Expected outcome: Because Cue A already predicts the outcome, the later presentation supplies less unexpected information than it would have supplied without the earlier conditioning.
Result for Cue B: Kamin blocking describes the case in which the new Cue B is learned weakly or not as expected because prior conditioning has changed the conditions for later learning.
Prior learning about Cue A changes the outcome of later learning about Cue B.
What do you think happens?
After Cue A has already been conditioned, what would the blocking account predict about learning for a new Cue B presented together with Cue A?
Reveal answer
Answer: Cue B may receive little or no additional learning.
Kamin blocking is defined in the source as a case in which prior conditioning affects subsequent learning. The previously conditioned cue makes the outcome less surprising when the compound is presented.
Surprise in Rescorla-Wagner
The Rescorla-Wagner model links learning to surprise. Learning occurs when animals are surprised, so the amount of new learning depends on the difference between what the learner expects and what occurs.
This idea gives blocking a learning-based explanation. Before the compound stage, Cue A has already acquired predictive significance. When Cue A and Cue B appear together, the outcome is not as unexpected as it would have been for a learner with no prior conditioning. Since the outcome is already prepared for, the new Cue B may contribute little or no additional learning.
To analyze a Rescorla-Wagner explanation, ask two questions: What does the learner already predict, and how much unexpected information remains when the outcome occurs?
Alternative Conditioning Models
The Rescorla-Wagner model is one model among several used to study classical conditioning. The source lists models associated with Klopf, Grossberg, Mackintosh, Moore and Stickney, Pearce and Hall, and Courville, Daw, and Touretzky. This broader list matters because it places Rescorla and Wagner within a history of animal learning theory rather than treating their model as the only available explanation.
| Model or model family named in the source | What can be stated from the source pack |
|---|---|
| Klopf | An alternative model associated with the study of classical conditioning |
| Grossberg | An alternative model associated with the study of classical conditioning |
| Mackintosh | A model associated with the broader history of animal learning theory |
| Moore and Stickney | An alternative model associated with the study of classical conditioning |
| Pearce and Hall | An alternative model associated with the study of classical conditioning |
| Courville, Daw, and Touretzky | An alternative model associated with the study of classical conditioning |
The source identifies these alternatives but does not specify their mechanisms in the supplied material.
Temporal Differences in a Conditioning Episode
The TD model places time and changes in prediction at the center of its account of classical conditioning. A temporal difference concerns the relationship between predictions at different moments in a conditioning episode.
Imagine following one conditioning episode from its beginning to its end. At the beginning, the learner has one set of expectations. As the episode unfolds, later events provide new information. The TD perspective compares these successive points rather than treating the whole episode as one undivided event. Learning is therefore connected to changes in predictive information over time.
Reading one episode as temporal change
Explain a simple conditioning episode from a TD perspective without treating the entire episode as one undifferentiated event.
Start with the beginning: Describe the learner's expectations at the beginning of the episode.
Advance to the cue: When the cue appears, it provides information at a particular moment. The model considers how this changes the predictive situation.
Advance to later information: As the episode continues toward the outcome, the learner receives further information and can compare the prediction at this moment with the prediction at an earlier moment.
Interpret the difference: The learning-relevant quantity is the time-sensitive change in predictive information, not merely the fact that the cue and outcome occurred in the same episode.
The TD account organizes conditioning around successive predictions and the differences between them.
A temporal difference is not simply a label for the final outcome. It identifies a time-sensitive change in predictive information that the model uses to organize learning.
Three Learning Rules in One History
The TD model developed within a larger history of learning rules. Sutton and Barto's 1981 work recognized a near identity between the Rescorla-Wagner model and the Least-Mean-Square learning rule, also called the LMS or Widrow-Hoff rule. Widrow and Hoff introduced the LMS rule in 1960. This connection places the TD model alongside an established error-based learning-rule tradition rather than presenting it as an unrelated account of conditioning.
| Item | Historical point |
|---|---|
| LMS or Widrow-Hoff rule | Introduced by Widrow and Hoff in 1960 |
| Rescorla-Wagner and LMS | Sutton and Barto's 1981 work recognized a near identity between them |
| Early TD model | An early version appeared in Sutton and Barto in 1981 |
| TD algorithm | Sutton later developed the algorithm in 1984 and 1988 |
| TD model presentations | The model was revised and presented in Sutton and Barto in 1987, with a more complete presentation in 1990 |
The source describes a gradual development rather than one finished theory appearing in a single publication.
The development was gradual. The early version of the TD model made a specific prediction: temporal primacy overrides blocking. The source reports that Kehoe, Scheurs, and Graham later showed this effect in the rabbit nictitating membrane preparation. This is an example of a model being evaluated against an experimental result, not merely a restatement that timing exists.
Why Representation Changes the Evaluation
The TD model does not depend only on its learning rule. The way a stimulus is represented can also affect how the model performs. Later work examined different stimulus representations and their possible neural implementations in relation to response timing and response topography.
A microstimulus representation represents a stimulus through time-specific microstimuli across an episode. The source emphasizes that this representation should not be assigned one fixed physical form automatically; it is a way of examining how stimulus information is encoded across time.
A temporal model needs a way to encode stimulus information across the episode. If the representation preserves useful distinctions between early and later moments, the model has a richer basis for comparing successive predictions. If those distinctions are not represented, the model has less information with which to organize time-sensitive changes.
Ludvig, Sutton, and Kehoe introduced the microstimulus representation in 2008. In 2012, they evaluated the TD model on previously unexplored classical-conditioning tasks and examined the influence of several stimulus representations, including the microstimulus representation. The important lesson is that evaluating a temporal learning rule also requires asking how the stimulus is represented.
Same stimulus, different representational question
Why might a TD evaluation ask whether a stimulus is represented as one undifferentiated event or through time-specific microstimuli?
Identify the TD requirement: The TD model compares predictive information at different moments, so it needs stimulus information that can be related to those moments.
Consider one undivided representation: If the stimulus is treated without useful time-specific distinctions, the model has a less detailed basis for comparing early and later predictions.
Consider microstimuli: A microstimulus representation examines the stimulus through a sequence of time-specific components, allowing the evaluation to ask which components support which predictions.
Draw the evaluation lesson: Differences in performance may reflect not only the learning rule but also the representation supplied to that rule.
Stimulus representation is part of the explanation being tested, not merely an implementation detail outside the model.
Common Reasoning Errors
Treating conditioning trials as independent events
Kamin blocking specifically depends on prior conditioning changing the conditions for later learning.
Fix:
Trace the learning history before interpreting what happens when the new stimulus is introduced.Defining learning as simple stimulus-outcome co-occurrence
The Rescorla-Wagner account emphasizes whether the outcome is surprising given what is already predicted.
Fix:
Ask what the learner already expects before deciding how much new learning remains.Treating the TD model as a model of only the final outcome
The TD model focuses on changes in predictive information across different moments.
Fix:
Compare neighboring moments and ask how the prediction changes as new information becomes available.Assuming the TD learning rule is sufficient without specifying a representation
The source states that stimulus representation, including microstimulus representation, affects how the TD model performs.
Fix:
Treat representation as part of the evaluation of the model.Assuming reinforcement has one universal historical meaning
The source describes meaningful differences among these traditions.
Fix:
Identify the historical tradition and its definition before comparing claims.Inventing mechanisms for every alternative model named in the source
The source lists these alternatives but does not provide their detailed mechanisms here.
Fix:
Use the supplied material only to identify them as part of the broader history of conditioning models.
Practice the Model Trace
A learner first experiences Cue A in a conditioning arrangement. Later, Cue A is presented together with a new Cue B. Write a short explanation of the result using both of these perspectives: first, explain why Cue B may receive little learning under the Rescorla-Wagner emphasis on surprise; second, explain what a TD analysis would examine across the episode.
Hints
- Begin with the learner's prior conditioning history for Cue A.
- For the Rescorla-Wagner perspective, ask whether the outcome is already expected.
- For the TD perspective, identify successive moments and the changes in predictive information between them.
- Mention why the representation of the stimulus may matter for the TD analysis.
A complete comparison
Compare the Rescorla-Wagner and TD explanations of the same two-stage conditioning situation.
State the prior learning: Cue A was conditioned in the first stage, so the learner enters the second stage with an established expectation associated with Cue A.
Apply the Rescorla-Wagner emphasis: When Cue A and Cue B appear together, the outcome is less surprising because Cue A already provides predictive information. Therefore Cue B may receive little or no additional learning.
Apply the TD emphasis: Follow the episode across time and compare predictions at successive moments instead of treating the compound and outcome as one undivided event.
Add representation: Ask how the stimulus is encoded across the episode. A time-sensitive representation, such as the microstimulus representation examined in later work, gives the model a basis for evaluating predictions at different moments.
The Rescorla-Wagner account emphasizes surprise and prior expectation, while the TD account emphasizes time-dependent changes in prediction and the representation that makes those changes available.
Key Takeaways
- The meaning of reinforcement changed across historical traditions in animal learning, so the term must be interpreted in context.
- Kamin blocking demonstrates that prior conditioning can affect later learning about a newly added stimulus.
- The Rescorla-Wagner model links learning with surprise: when an outcome is already expected, additional learning may be small.
- The TD model explains conditioning through changes in predictive information across time and developed within a history connected to the Rescorla-Wagner model and the LMS learning rule.
- Stimulus representation matters for TD evaluations, and microstimulus representation provides a way to examine stimulus information through time-specific components.
Key Takeaways
- Prior conditioning changes the conditions for later learning, as shown by Kamin blocking.
- The Rescorla-Wagner model emphasizes surprise and explains why an already expected outcome may produce little new learning.
- The TD model compares predictions across moments in a conditioning episode rather than treating the episode as a single undivided event.
- The TD model has a historical relationship with the Rescorla-Wagner model and the LMS or Widrow-Hoff learning rule.
- Stimulus representation, including microstimulus representation, affects how the TD model can be evaluated on conditioning tasks.