Concepts / Animal Training

Animal Training

Shaping changes the reward rule gradually rather than demanding the final behavior at once.

  • Programming

Why the Final Behavior Is Not Enough

Suppose an animal or learning agent has not yet learned a complicated behavior. If training rewards only the complete behavior, the learner may have no practical route toward discovering it. Shaping addresses this problem by rewarding behavior that moves toward the target and then gradually changing what counts as successful.

Shaping does not demand the final behavior immediately. It changes the reward rule gradually.

What do you think happens?

What is the central change during shaping?

  • The learner is required to perform the complete behavior from the beginning
  • The behavior that earns a reward becomes a closer approximation of the desired behavior
  • The learner stops responding to events
Reveal answer

Answer: The behavior that earns a reward becomes a closer approximation of the desired behavior

Each training stage rewards a behavior that is closer to the desired behavior than the behavior rewarded at the preceding stage.

Successive Approximations

A successive approximation is a behavior that moves the learner closer to the desired behavior. In shaping, the trainer or training system first rewards an attainable behavior related to the target. After that behavior becomes the current success condition, the reward contingency changes so that a closer behavior is rewarded. Repeating this process guides the learner through a sequence of increasingly close approximations.

reward rule changesreward rule changesInitial behaviorfirst approximationCloser behaviornext approximationDesired behaviortarget
What changes from one training step to the next as the learner moves from a simple behavior toward the desired behavior?

The important change is not merely that time passes. The condition for receiving reward is adjusted from stage to stage. This makes the sequence directional: each rewarded behavior is selected because it is a closer approximation of the desired behavior.

A Training Sequence

Moving a Learner Toward a Target

Illustrate how a training rule can move from rewarding a simple behavior to rewarding a desired behavior.

Stage 1: Reward a simple behavior that moves in the direction of the target. The complete desired behavior is not required yet.

Stage 2: Change the contingency so that a behavior closer to the target receives the reward.

Stage 3: Change the contingency again so that the desired behavior itself, or a still closer approximation, receives the reward.

The learner is guided through successive approximations rather than being required to discover the complete behavior at once.

This is a conceptual example rather than a particular animal-training protocol. Its purpose is to show the structure of shaping: an initial reward condition is replaced by a closer condition, which is later replaced by another closer condition.

Changing Reward Contingencies

A reward contingency is the relationship between a behavior and the reward that follows it. In shaping, this relationship changes over successive training stages so that increasingly close behaviors receive the reward.

contingency changescontingency changes againSimple behaviorearns rewardCloser behaviorearns rewardDesired behaviorlater success condition
How does the behavior that earns a reward change over successive stages of training?

Animals and Learning Agents

Shaping is important in animal training, but the same staged structure can also train reinforcement learning agents. Reinforcement learning includes a control aspect based on learning through trial and error. Shaping organizes that process by using changing reward contingencies instead of requiring the agent to discover the complete desired behavior at once.

behavior producesguidesaction producesguidesAnimalbehaviorLearning agentactionRewardchanging contingencyRewardchanging contingencyCloser behaviorsuccessive approximation
How does information about an action and its reward move through training for an animal compared with a reinforcement learning agent?

The connection is about the training structure, not about claiming that an animal and an artificial agent learn in every respect in the same way. In both cases described here, the useful pattern is staged adjustment of what receives reward.

Shaping connects animal training and reinforcement learning through trial and error, consequences, and progressively adjusted reward conditions.

Motivational State

Training does not occur independently of the learner's motivational state. Motivational state affects how an animal responds to events, including whether it approaches or avoids them. Therefore, the same event cannot be interpreted without considering the animal's current motivational state.

influences response tomay producemay produceMotivational statecurrent conditionEventexperienced by animalApproachpossible responseAvoidancepossible response
How does an animal's motivational state change which events it approaches and which it avoids?

This point adds an important qualification to a simple reward-centered description of training. What an animal approaches or avoids is influenced by its motivational state, so behavior and consequences must be considered together with the learner's current state.

Mistakes About Shaping

  • Treating shaping as rewarding only the final behavior

    Shaping is designed to provide a route toward the target by rewarding behaviors that are progressively closer to it.

    Fix: Describe the training as a sequence in which the rewarded behavior changes from an initial approximation to closer approximations.

  • Assuming that the reward condition stays fixed

    A fixed reward condition does not express the staged adjustment central to shaping.

    Fix: Identify how the reward contingency changes from one stage to the next.

  • Separating shaping completely from reinforcement learning

    The method can also be used to train reinforcement learning agents.

    Fix: Connect both cases through trial and error and changing reward contingencies.

  • Ignoring motivational state

    Motivational state affects how the animal responds to events.

    Fix: Include motivational state when analyzing whether an event is approached or avoided.

Apply the Mechanism

MEDIUM

A learner has not yet performed a complicated target behavior. Describe a three-stage shaping plan in words. For each stage, state what kind of behavior receives the reward and explain why the next stage must use a closer approximation.

Hints
  • Begin with a behavior that moves toward the target rather than requiring the complete target.
  • Make the rewarded behavior closer to the desired behavior at each later stage.
  • Use the phrase changing reward contingency to explain what changes.
MEDIUM

Compare animal training with training a reinforcement learning agent. Identify the shared structure, then explain why motivational state is an additional consideration when discussing an animal's approach or avoidance of events.

Hints
  • The shared structure involves trial and error and changing what receives reward.
  • Motivational state affects how an animal responds to events.

Key Takeaways

  1. Shaping changes the reward rule gradually instead of demanding the final behavior immediately.
  2. Successive approximations are behaviors that move progressively closer to the desired behavior.
  3. The same staged adjustment of reward contingencies can organize animal training and the training of reinforcement learning agents.
  4. Motivational state influences whether an animal approaches or avoids an event.
  5. Shaping belongs to the broader picture of learning through trial and error, in which consequences influence behavior.

Key Takeaways

  • Shaping uses progressively changing reward contingencies.
  • Each stage rewards a closer approximation of the desired behavior.
  • The method applies to animal training and can also train reinforcement learning agents.
  • Motivational state affects an animal's approach to or avoidance of events.