Concepts / Reinforcement Learning and Animal Learning

Reinforcement Learning and Animal Learning

The defining question is whether reinforcement depends on the animal's behavior.

  • Programming

The Contingency Question

A useful way to analyze an animal-learning situation is to ask one question: does the reinforcing event depend on what the animal does? If the answer is yes, the situation is instrumental conditioning. The animal's behavior is part of the condition for receiving the reinforcing stimulus. If the reinforcing stimulus does not depend on the animal's behavior, the situation belongs to classical, or Pavlovian, conditioning instead.

involvesdetermines delivery ofpresentsdoes not depend on behaviorInstrumentalconditioningAnimal behaviorContingentreinforcementClassicalconditioningReinforcing eventBehavior-independentreinforcement
Does the animal's behavior determine whether the reinforcing event is delivered?

The decisive feature is not simply that reinforcement occurs. It is whether the animal's behavior is required for the reinforcing event to occur.

From Behavior to Control

Instrumental conditioning connects behavior with contingent consequences. This gives it a control-oriented character: the learner is not only acquiring information about what predicts a reinforcing event, but is also involved in obtaining that event through behavior. In this sense, reinforcement learning concerns the relationship between what the learner does and what follows.

producesis part ofdetermines access tosupportsLearnerBehaviorContingentconsequenceReinforcing stimulusControl
What changes when reinforcement is contingent on the animal performing a particular behavior?

Classifying a Learning Situation

Suppose a learner receives a reinforcing event only after performing a particular response. Is this situation instrumental or classical conditioning?

Identify the candidate behavior: The learner performs a response before the reinforcing event occurs.

Check the contingency: The reinforcing event depends on that response, so behavior is part of the condition for receiving reinforcement.

Classify the situation: Because reinforcement is behavior-dependent, the situation is instrumental conditioning.

The situation is instrumental conditioning. If the reinforcing event had occurred independently of the learner's behavior, it would instead belong to classical, or Pavlovian, conditioning.

Trial and Error

Control requires the learner to try behavior and use what happens to guide later behavior. This is the role of trial and error in reinforcement learning. Different actions can produce different outcomes, and those outcomes provide information for future action choices. Thorndike's experiments with cats are associated with this tradition and with learning by trial and error.

leads toproduceguidesinfluencesCurrent situationDifferent actionsAction outcomesAdjusted futurechoicesLater behavior
How do different actions produce outcomes that guide later action choices?

Trial and error does not necessarily mean a completely blind search. Trials can be generated using innate knowledge and knowledge acquired earlier, provided that some exploration remains. The important point is that the learner tests behavior and uses the resulting consequences to guide later behavior.

What do you think happens?

A learner has already acquired some knowledge about which actions tend to produce reinforcement. Does trial and error require the learner to ignore that knowledge?

  • Yes. Every action must be selected blindly.
  • No. Earlier or innate knowledge can guide trials, although some exploration remains.
  • Yes. Only the most familiar action can be selected.
Reveal answer

Answer: No. Earlier or innate knowledge can guide trials, although some exploration remains.

The source describes exploration as compatible with knowledge already available to the learner. Trial and error remains important because behavior is still tried and consequences are used to guide later choices.

Shaping Toward a Target

A desired behavior may be too difficult to obtain immediately. Shaping addresses this problem by progressively altering the reward contingencies. The learner is trained through successive approximations: early contingencies support behavior that is closer to the target, and later contingencies support behavior that more closely matches the desired result.

approachesapproachessupportssupportsApproximatebehaviorCloser behaviorDesired behaviorEarly contingencyTarget contingency
How does the rewarded behavior change step by step as reinforcement shifts from an approximate response toward the desired behavior?

Reading a Shaping Sequence

A trainer wants a learner to perform a target behavior that is initially too difficult to obtain. How does shaping alter the training process?

Begin with an approximation: The initial reward contingency supports behavior that is closer to the target than the learner's starting behavior.

Shift the contingency: As the learner produces closer behavior, reinforcement is reserved for a response that more closely matches the desired result.

Reach the target: The progression ends with a contingency that supports behavior matching the desired target.

Shaping is a progressive change in what behavior is reinforced. Its central idea is not simply to provide more reinforcement, but to move the reinforcement contingency toward the desired behavior.

Prediction and Control

TD algorithms belong on the prediction side of the reinforcement-learning comparison. The source connects TD learning with classical conditioning because TD algorithms learn to predict. This makes TD methods different from control methods: prediction concerns learning what to expect, whereas control concerns using behavior and its consequences to guide later behavior.

learns fromsupportsinvolvesobtainsTD algorithmsSuccessiveexperiencesPredictionReinforcement-learningcontrolBehaviorContingentconsequences
How does TD learning as prediction differ from choosing actions to control future rewards?
AspectTD algorithmsControl-oriented reinforcement learning
Primary roleLearn to predictUse behavior and consequences to guide later behavior
Analogy in animal learningClassical conditioningInstrumental conditioning
Central questionWhat should be predicted from successive experiences?Which behavior helps obtain reinforcing consequences?

The TD model also includes the temporal dimension of events within individual trials, generalizes the Rescorla-Wagner model, and provides an account of second-order conditioning, in which predictors of reinforcing stimuli become reinforcing themselves. These details reinforce the prediction connection; they do not turn TD algorithms into control methods.

Common Classification Errors

  • Treating every reinforcement-learning process as classical conditioning.

    The reinforcing event depends on the learner's behavior, which is the defining feature of instrumental conditioning.

    Fix: Check the behavior-reinforcement contingency before choosing the classification.

  • Assuming that trial and error means completely blind exploration.

    Exploration can be generated using innate knowledge and knowledge acquired earlier.

    Fix: Allow prior knowledge to guide trials while retaining some exploration.

  • Defining shaping as simply increasing the amount of reinforcement.

    Shaping changes which behavior is reinforced as the learner approaches the target.

    Fix: Describe the progressive movement of the reward contingency from an approximation toward the desired behavior.

  • Calling TD algorithms control methods because they are discussed in reinforcement learning.

    The source places TD algorithms on the prediction side and connects them with classical conditioning.

    Fix: Describe TD algorithms as methods for learning predictions from successive experiences.

Check Your Understanding

MEDIUM

For each situation, identify the most relevant concept and justify your choice: instrumental conditioning, classical conditioning, trial and error, shaping, or TD prediction. Situation A: a reinforcing event is delivered only when the learner performs a particular behavior. Situation B: a learner tries behaviors, observes their consequences, and uses those consequences to guide later behavior. Situation C: the rewarded response is progressively changed so that it approaches a desired target. Situation D: a method learns predictions from successive experiences without being described as a method for selecting behavior to control future consequences.

Hints
  • For Situation A, ask whether reinforcement depends on behavior.
  • For Situation B, focus on how outcomes influence later choices.
  • For Situation C, focus on the changing reward contingency.
  • For Situation D, distinguish prediction from control.
  1. Use the behavior-contingency question first: behavior-dependent reinforcement indicates instrumental conditioning, while behavior-independent reinforcement indicates classical or Pavlovian conditioning. Instrumental conditioning has a control-oriented character because the learner's behavior helps obtain the consequence. Trial and error connects action outcomes with later choices, and it can be guided by prior knowledge rather than being completely blind. Shaping progressively changes which behavior is reinforced so that behavior approaches a desired target. TD algorithms belong primarily to prediction and correspond to classical conditioning in the reinforcement-learning analogy, whereas control methods concern behavior and its consequences.

Key Takeaways

  • Instrumental conditioning is defined by reinforcement that depends on the animal's behavior.
  • Classical or Pavlovian conditioning involves a reinforcing stimulus that does not depend on the animal's behavior.
  • Trial and error supports control by allowing outcomes of behavior to guide later behavior, with exploration potentially guided by prior knowledge.
  • Shaping progressively changes reward contingencies from approximate responses toward a desired behavior.
  • TD algorithms are prediction methods associated with classical conditioning in the analogy, not control methods for selecting behavior.