Animal Training
Shaping changes the reward rule gradually rather than demanding the final behavior at once.
Why the Final Behavior Is Not Enough
Suppose an animal or learning agent has not yet learned a complicated behavior. If training rewards only the complete behavior, the learner may have no practical route toward discovering it. Shaping addresses this problem by rewarding behavior that moves toward the target and then gradually changing what counts as successful.
Shaping does not demand the final behavior immediately. It changes the reward rule gradually.
What do you think happens?
What is the central change during shaping?
Reveal answer
Answer: The behavior that earns a reward becomes a closer approximation of the desired behavior
Each training stage rewards a behavior that is closer to the desired behavior than the behavior rewarded at the preceding stage.
Successive Approximations
A successive approximation is a behavior that moves the learner closer to the desired behavior. In shaping, the trainer or training system first rewards an attainable behavior related to the target. After that behavior becomes the current success condition, the reward contingency changes so that a closer behavior is rewarded. Repeating this process guides the learner through a sequence of increasingly close approximations.
The important change is not merely that time passes. The condition for receiving reward is adjusted from stage to stage. This makes the sequence directional: each rewarded behavior is selected because it is a closer approximation of the desired behavior.
A Training Sequence
Moving a Learner Toward a Target
Illustrate how a training rule can move from rewarding a simple behavior to rewarding a desired behavior.
Stage 1: Reward a simple behavior that moves in the direction of the target. The complete desired behavior is not required yet.
Stage 2: Change the contingency so that a behavior closer to the target receives the reward.
Stage 3: Change the contingency again so that the desired behavior itself, or a still closer approximation, receives the reward.
The learner is guided through successive approximations rather than being required to discover the complete behavior at once.
This is a conceptual example rather than a particular animal-training protocol. Its purpose is to show the structure of shaping: an initial reward condition is replaced by a closer condition, which is later replaced by another closer condition.
Changing Reward Contingencies
A reward contingency is the relationship between a behavior and the reward that follows it. In shaping, this relationship changes over successive training stages so that increasingly close behaviors receive the reward.
Animals and Learning Agents
Shaping is important in animal training, but the same staged structure can also train reinforcement learning agents. Reinforcement learning includes a control aspect based on learning through trial and error. Shaping organizes that process by using changing reward contingencies instead of requiring the agent to discover the complete desired behavior at once.
The connection is about the training structure, not about claiming that an animal and an artificial agent learn in every respect in the same way. In both cases described here, the useful pattern is staged adjustment of what receives reward.
Shaping connects animal training and reinforcement learning through trial and error, consequences, and progressively adjusted reward conditions.
Motivational State
Training does not occur independently of the learner's motivational state. Motivational state affects how an animal responds to events, including whether it approaches or avoids them. Therefore, the same event cannot be interpreted without considering the animal's current motivational state.
This point adds an important qualification to a simple reward-centered description of training. What an animal approaches or avoids is influenced by its motivational state, so behavior and consequences must be considered together with the learner's current state.
Mistakes About Shaping
Treating shaping as rewarding only the final behavior
Shaping is designed to provide a route toward the target by rewarding behaviors that are progressively closer to it.
Fix:
Describe the training as a sequence in which the rewarded behavior changes from an initial approximation to closer approximations.Assuming that the reward condition stays fixed
A fixed reward condition does not express the staged adjustment central to shaping.
Fix:
Identify how the reward contingency changes from one stage to the next.Separating shaping completely from reinforcement learning
The method can also be used to train reinforcement learning agents.
Fix:
Connect both cases through trial and error and changing reward contingencies.Ignoring motivational state
Motivational state affects how the animal responds to events.
Fix:
Include motivational state when analyzing whether an event is approached or avoided.
Apply the Mechanism
A learner has not yet performed a complicated target behavior. Describe a three-stage shaping plan in words. For each stage, state what kind of behavior receives the reward and explain why the next stage must use a closer approximation.
Hints
- Begin with a behavior that moves toward the target rather than requiring the complete target.
- Make the rewarded behavior closer to the desired behavior at each later stage.
- Use the phrase changing reward contingency to explain what changes.
Compare animal training with training a reinforcement learning agent. Identify the shared structure, then explain why motivational state is an additional consideration when discussing an animal's approach or avoidance of events.
Hints
- The shared structure involves trial and error and changing what receives reward.
- Motivational state affects how an animal responds to events.
Key Takeaways
- Shaping changes the reward rule gradually instead of demanding the final behavior immediately.
- Successive approximations are behaviors that move progressively closer to the desired behavior.
- The same staged adjustment of reward contingencies can organize animal training and the training of reinforcement learning agents.
- Motivational state influences whether an animal approaches or avoids an event.
- Shaping belongs to the broader picture of learning through trial and error, in which consequences influence behavior.
Key Takeaways
- Shaping uses progressively changing reward contingencies.
- Each stage rewards a closer approximation of the desired behavior.
- The method applies to animal training and can also train reinforcement learning agents.
- Motivational state affects an animal's approach to or avoidance of events.