Concepts / Classical Conditioning and Prediction

Classical Conditioning and Prediction

The defining question is whether reinforcement depends on the animal's behavior.

  • Programming

One Contingency Question

A useful way to classify a learning situation is to ask one question: does the reinforcing event depend on what the animal does? If the answer is yes, the situation is instrumental conditioning. If the reinforcing stimulus does not depend on the animal's behavior, the situation is classical, or Pavlovian, conditioning. This difference separates learning that helps control consequences from learning that primarily concerns prediction.

InstrumentalconditioningReinforcement depends onbehaviorClassicalconditioningReinforcement does notdepend on behavior
Does reinforcement occur independently of behavior, or does the animal's behavior determine whether reinforcement is delivered?

Tracing the Reinforcement Contingency

Classifying two learning situations

Determine whether each situation is instrumental or classical conditioning.

Situation A: An animal performs a behavior, and that behavior is part of the condition for receiving the reinforcing stimulus. Because reinforcement depends on the behavior, this is instrumental conditioning.

Situation B: A reinforcing stimulus is delivered without depending on what the animal does. Because the animal's behavior is not part of the condition for receiving the stimulus, this is classical, or Pavlovian, conditioning.

Decision rule: Do not classify the situation by asking only whether a reward appears. Classify it by checking whether the learner's behavior is required for the reinforcing event.

Behavior-dependent reinforcement indicates instrumental conditioning; behavior-independent reinforcement indicates classical conditioning.

Instrumental conditioning connects behavior with contingent consequences. The behavior is involved in obtaining the reinforcing event, so the learning situation has a control-oriented character. Classical conditioning differs in the defining contingency: the reinforcing stimulus does not depend on the animal's behavior. In that case, the important relationship is not that a behavior produces the event, but that events can provide information about what is predicted.

supportscompared with outcomeprovidesupdatesnot identical toCue or current eventsInformation about a futurerewardReward predictionCurrent expected valueReinforcing eventWhat occursPrediction errorDifference betweenprediction and outcomeUpdated predictionRevised expected valueBehavior choiceA control question, notonly a prediction
How does a prediction error update the expected value of a future reward, and how is this different from choosing an action to control the reward?

Trial and Error as Control

Control requires the learner to try behavior and use what happens to guide later behavior. This is the role of trial and error in reinforcement learning. A trial is not merely an observation of a prediction; it includes behavior whose consequences can influence what the learner does later. Thorndike's experiments with cats are associated with this tradition and with learning by trial and error.

producesguidescontinues the loopBehaviorTry an actionConsequenceWhat happens after thebehaviorLater behaviorUse feedback to guideanother trial
How does an agent's action lead to feedback that changes which action it is likely to choose next?

Trial and error does not require a completely blind search. Trials can be generated using innate knowledge and knowledge acquired earlier, provided that some exploration remains. The essential control loop is that the learner tries behavior, observes what happens, and uses the feedback to guide later behavior.

Generated example: Imagine an agent choosing among behaviors while learning a task. If one behavior is followed by a reinforcing consequence, the agent can use that result when deciding what to try later. The example is instrumental because the behavior is part of the condition for obtaining the consequence; it is not merely a cue whose associated event is being predicted.

Shaping Through Successive Approximations

A desired behavior may be too difficult to obtain immediately. Shaping addresses this problem by progressively altering the reward contingencies. Early contingencies support behavior that is closer to the target; later contingencies support behavior that more closely matches the desired result. The learner therefore moves through successive approximations rather than needing to produce the final behavior at once.

progressively adjusts reinforcementCloser behaviorEarly contingencyDesired behaviorLater contingency
How do successive approximations to a desired behavior change the behavior that receives reinforcement?

Following the logic of shaping

Explain how shaping can train a behavior that is initially too difficult to obtain directly.

Begin near the target: Arrange the initial reward contingency so that behavior closer to the desired behavior receives reinforcement.

Change the contingency: As the learner produces behavior closer to the target, alter what behavior is reinforced so that the requirement moves nearer to the desired result.

Reach the target: Continue the progressive adjustment until the reinforced behavior more closely matches the desired behavior.

Shaping is a change in what behavior receives reinforcement, carried out through successive approximations toward a desired behavior.

TD Algorithms and Prediction

A temporal-difference algorithm belongs on the prediction side of the comparison. The source connects TD learning with classical conditioning because TD algorithms learn to predict. A prediction error can be used to revise the expected value of a future reinforcing event. This makes TD learning relevant to the changing relationship between predictors and reinforcing events, rather than making it a control method by itself.

The TD model also includes the temporal dimension of events within individual trials, generalizes the Rescorla-Wagner model, and provides an account of second-order conditioning. In second-order conditioning, predictors of reinforcing stimuli become reinforcing themselves. These points explain why TD algorithms are associated with classical conditioning and prediction: they model how predictions change across temporally arranged events.

includesincludesPredictionWhat reinforcing event isexpectedTD algorithmsLearn to predictControlWhich behavior obtainsreinforcementTrial and errorTry behavior and usefeedback
What is the difference between learning what reward is predicted by a cue and learning which behavior will produce a reward?

Check the Contingency

MEDIUM

For each description, identify the central learning problem: prediction or control. Then explain whether it is more closely associated with classical conditioning or instrumental conditioning. Finally, state whether shaping could apply and why.

Hints
  • First ask whether reinforcement depends on the learner's behavior.
  • For prediction, focus on what reinforcing event is expected.
  • For control, focus on trying behavior and using feedback to guide later behavior.
  • For shaping, look for progressive changes in the behavior that receives reinforcement.

What do you think happens?

An algorithm learns how the value of a future reinforcing event changes when its prediction differs from what occurs. Is this primarily a prediction mechanism or a control mechanism?

  • Prediction mechanism
  • Control mechanism
Reveal answer

Answer: Prediction mechanism

TD algorithms are associated with prediction and classical conditioning in the reinforcement-learning analogy. Control additionally requires trying behavior and using consequences to guide later behavior.

The Working Distinction

  1. The defining question is whether reinforcement depends on the animal's behavior.
  2. Instrumental conditioning is behavior-dependent and therefore has a control-oriented character.
  3. Classical conditioning is behavior-independent in this comparison and is associated with prediction.
  4. Trial and error supplies control by linking tried behavior with feedback that guides later behavior.
  5. Shaping progressively changes reward contingencies through successive approximations toward a desired behavior.
  6. TD algorithms learn to predict and should not be confused with methods for choosing behavior to control reinforcement.

Key Takeaways

  • Classify conditioning by checking whether reinforcement depends on behavior.
  • Instrumental conditioning concerns behavior and its contingent consequences, whereas classical conditioning concerns prediction when the reinforcing stimulus does not depend on behavior.
  • Trial and error supports control because feedback from tried behavior guides later behavior.
  • Shaping progressively changes which behavior is reinforced so behavior approaches a desired target.
  • TD algorithms belong primarily to prediction and are connected with classical conditioning, not identical to control methods.