Classical Conditioning and Prediction
The defining question is whether reinforcement depends on the animal's behavior.
One Contingency Question
A useful way to classify a learning situation is to ask one question: does the reinforcing event depend on what the animal does? If the answer is yes, the situation is instrumental conditioning. If the reinforcing stimulus does not depend on the animal's behavior, the situation is classical, or Pavlovian, conditioning. This difference separates learning that helps control consequences from learning that primarily concerns prediction.
Tracing the Reinforcement Contingency
Classifying two learning situations
Determine whether each situation is instrumental or classical conditioning.
Situation A: An animal performs a behavior, and that behavior is part of the condition for receiving the reinforcing stimulus. Because reinforcement depends on the behavior, this is instrumental conditioning.
Situation B: A reinforcing stimulus is delivered without depending on what the animal does. Because the animal's behavior is not part of the condition for receiving the stimulus, this is classical, or Pavlovian, conditioning.
Decision rule: Do not classify the situation by asking only whether a reward appears. Classify it by checking whether the learner's behavior is required for the reinforcing event.
Behavior-dependent reinforcement indicates instrumental conditioning; behavior-independent reinforcement indicates classical conditioning.
Instrumental conditioning connects behavior with contingent consequences. The behavior is involved in obtaining the reinforcing event, so the learning situation has a control-oriented character. Classical conditioning differs in the defining contingency: the reinforcing stimulus does not depend on the animal's behavior. In that case, the important relationship is not that a behavior produces the event, but that events can provide information about what is predicted.
Trial and Error as Control
Control requires the learner to try behavior and use what happens to guide later behavior. This is the role of trial and error in reinforcement learning. A trial is not merely an observation of a prediction; it includes behavior whose consequences can influence what the learner does later. Thorndike's experiments with cats are associated with this tradition and with learning by trial and error.
Trial and error does not require a completely blind search. Trials can be generated using innate knowledge and knowledge acquired earlier, provided that some exploration remains. The essential control loop is that the learner tries behavior, observes what happens, and uses the feedback to guide later behavior.
Generated example: Imagine an agent choosing among behaviors while learning a task. If one behavior is followed by a reinforcing consequence, the agent can use that result when deciding what to try later. The example is instrumental because the behavior is part of the condition for obtaining the consequence; it is not merely a cue whose associated event is being predicted.
Shaping Through Successive Approximations
A desired behavior may be too difficult to obtain immediately. Shaping addresses this problem by progressively altering the reward contingencies. Early contingencies support behavior that is closer to the target; later contingencies support behavior that more closely matches the desired result. The learner therefore moves through successive approximations rather than needing to produce the final behavior at once.
Following the logic of shaping
Explain how shaping can train a behavior that is initially too difficult to obtain directly.
Begin near the target: Arrange the initial reward contingency so that behavior closer to the desired behavior receives reinforcement.
Change the contingency: As the learner produces behavior closer to the target, alter what behavior is reinforced so that the requirement moves nearer to the desired result.
Reach the target: Continue the progressive adjustment until the reinforced behavior more closely matches the desired behavior.
Shaping is a change in what behavior receives reinforcement, carried out through successive approximations toward a desired behavior.
TD Algorithms and Prediction
A temporal-difference algorithm belongs on the prediction side of the comparison. The source connects TD learning with classical conditioning because TD algorithms learn to predict. A prediction error can be used to revise the expected value of a future reinforcing event. This makes TD learning relevant to the changing relationship between predictors and reinforcing events, rather than making it a control method by itself.
The TD model also includes the temporal dimension of events within individual trials, generalizes the Rescorla-Wagner model, and provides an account of second-order conditioning. In second-order conditioning, predictors of reinforcing stimuli become reinforcing themselves. These points explain why TD algorithms are associated with classical conditioning and prediction: they model how predictions change across temporally arranged events.
Check the Contingency
For each description, identify the central learning problem: prediction or control. Then explain whether it is more closely associated with classical conditioning or instrumental conditioning. Finally, state whether shaping could apply and why.
Hints
- First ask whether reinforcement depends on the learner's behavior.
- For prediction, focus on what reinforcing event is expected.
- For control, focus on trying behavior and using feedback to guide later behavior.
- For shaping, look for progressive changes in the behavior that receives reinforcement.
What do you think happens?
An algorithm learns how the value of a future reinforcing event changes when its prediction differs from what occurs. Is this primarily a prediction mechanism or a control mechanism?
Reveal answer
Answer: Prediction mechanism
TD algorithms are associated with prediction and classical conditioning in the reinforcement-learning analogy. Control additionally requires trying behavior and using consequences to guide later behavior.
The Working Distinction
- The defining question is whether reinforcement depends on the animal's behavior.
- Instrumental conditioning is behavior-dependent and therefore has a control-oriented character.
- Classical conditioning is behavior-independent in this comparison and is associated with prediction.
- Trial and error supplies control by linking tried behavior with feedback that guides later behavior.
- Shaping progressively changes reward contingencies through successive approximations toward a desired behavior.
- TD algorithms learn to predict and should not be confused with methods for choosing behavior to control reinforcement.
Key Takeaways
- Classify conditioning by checking whether reinforcement depends on behavior.
- Instrumental conditioning concerns behavior and its contingent consequences, whereas classical conditioning concerns prediction when the reinforcing stimulus does not depend on behavior.
- Trial and error supports control because feedback from tried behavior guides later behavior.
- Shaping progressively changes which behavior is reinforced so behavior approaches a desired target.
- TD algorithms belong primarily to prediction and are connected with classical conditioning, not identical to control methods.