Concepts / Goal Gradients and Delayed Reinforcement

Goal Gradients and Delayed Reinforcement

A delayed primary reward does not have to be the only influence operating during the delay.

  • Programming

Why Delay Creates a Learning Problem

In instrumental conditioning, an action and its reward do not always occur together. When the primary reward follows the action only after an interval, the reward has a longer distance to bridge before it can influence learning. Hull's account addresses this problem by proposing that the primary reward does not have to be the only influence operating during the delay.

The important question is not simply whether a delay exists. It is whether something during that delay can provide a more immediate reinforcing influence.

Tracing the Action-to-Reward Interval

Start with the sequence proposed in Hull's account. An action occurs first. A delay follows. The primary reward arrives at the end of that delay. If stimuli regularly appear during the interval, those stimuli can favor the development of secondary reinforcement. The learner is therefore not influenced only at the final point; an additional reinforcing influence can operate during the interval.

followed bycontainscan favor development ofends withActioninstrumental responseDelayintervalRegular stimuliduring the delaySecondaryreinforcementmore immediate influencePrimary rewardat the end
What happens between an action and a later primary reward when regularly occurring stimuli are present?

Secondary Reinforcement as a Bridge

In the context of delayed reinforcement, secondary reinforcement is the reinforcing influence that develops from stimuli regularly occurring during the delay. It supplements the delayed primary reward by providing a more immediate influence during the interval.

The source presents regularly occurring stimuli as conditions that can favor secondary reinforcement. The key feature is their position in the sequence: they occur after the action but before the primary reward. This gives them a possible role in connecting the earlier action with the later outcome.

is followed bycan favorsupportsalso influencesis separated from by delayActionearlier responseDelay stimulusregularly occurringSecondaryreinforcementimmediate influencePrimary rewardlater outcomeLearningsupported across delay
How is a stimulus during the delay related to the response and the later primary reward?

Extending the Goal Gradient

The source describes a goal gradient as something that can span time. Molar stimulus traces, understood in Hull's theory as the lingering presence of a stimulus, are part of the explanation. Hull's proposal is that secondary reinforcement works together with these traces, allowing the resulting gradient to extend beyond the period that stimulus traces alone would cover.

In practical terms, extending the goal gradient means that the influence associated with the goal can reach farther through the delay. Without secondary reinforcement, the primary reward is separated from the action by the entire interval. With secondary reinforcement, a more immediate influence appears within that interval, so learning need not rely only on the reward at the end.

followed byends withfollowed during delay byprecedesActioninitial responseActioninitial responseDelayno intermediate influenceDelay stimulusintermediate influencePrimary rewardfinal influencePrimary rewardfinal influence
How does secondary reinforcement change the pattern of influence across the delay compared with waiting only for the final reward?
delayeventually reachesdelay includescontinues towardActionbeginning of intervalDelayprimary reward stilldistantDelay stimulussecondary influencePrimary rewardend of intervalPrimary rewardlater outcome
How can secondary reinforcement alter the goal gradient as the learner moves through the delay toward the primary reward?

Supported and Obstructed Delays

Comparing Two Delays

Compare two situations in which an action is followed by a delayed primary reward. In one situation, regularly occurring stimuli during the delay support secondary reinforcement. In the other, conditions obstruct that secondary reinforcement.

Situation A: The delay includes regularly occurring stimuli that can favor secondary reinforcement. A more immediate reinforcing influence is therefore available during the interval.

Situation B: The conditions obstruct secondary reinforcement. The learner must rely more heavily on the primary reward at the end of the delay.

Prediction from Hull's account: Learning should decrease with increased delay less in Situation A than in Situation B, because the two delays provide different opportunities for secondary reinforcement.

The relevant comparison is not simply delay versus no delay. It is a delay with conditions favoring secondary reinforcement versus a delay in which secondary reinforcement is obstructed.

followed during delay byprecedesfollowed during delay byprecedesActionfollowed by delayActionfollowed by delayRegular stimulisecondary reinforcementsupportedObstructionsecondary reinforcementblockedPrimary rewardlater outcomePrimary rewardlater outcome
What differs when intermediate reinforcing stimuli are available during the delay versus blocked or absent?

Mistakes in Comparing Delays

  • Treating the delayed primary reward as the only influence during the interval.

    Hull's account proposes that regularly occurring stimuli during the delay can favor secondary reinforcement.

    Fix: Inspect the delay for stimuli that may provide a more immediate reinforcing influence.

  • Treating every stimulus during the delay as automatically reinforcing.

    The source says that regularly occurring stimuli can favor secondary reinforcement; it does not say that every stimulus necessarily becomes reinforcing.

    Fix: Describe the stimulus as a condition that can support or favor secondary reinforcement.

  • Comparing a delay only with no delay.

    The important comparison is between delays with different opportunities for secondary reinforcement.

    Fix: Compare a delay that supports secondary reinforcement with a delay in which secondary reinforcement is obstructed.

  • Confusing secondary reinforcement with the primary reward.

    Secondary reinforcement is the more immediate influence during the interval, whereas the primary reward remains the later outcome.

    Fix: Keep the intermediate reinforcing influence and the delayed primary reward conceptually separate.

  • Ignoring the role of stimulus traces.

    The source presents secondary reinforcement as working in conjunction with molar stimulus traces.

    Fix: Include both the lingering stimulus traces and the secondary reinforcement proposed to extend the gradient.

Apply the Distinction

MEDIUM

A response is followed by a delay and then by a primary reward. During the delay, regularly occurring stimuli are present in one condition but secondary reinforcement is obstructed in another. Explain which condition should show the smaller decrease in learning as the delay increases, and identify the mechanism that accounts for the difference.

Hints
  • Look for what occurs during the delay, not only for what occurs at its end.
  • Identify whether the conditions favor or obstruct secondary reinforcement.
  • Use the comparison between the two delays rather than comparing either delay with no delay.

Practice Answer

Which condition should show the smaller decrease in learning as the delay increases?

Identify the supported condition: The condition with regularly occurring stimuli provides an opportunity for secondary reinforcement during the delay.

Identify the obstructed condition: The other condition lacks that opportunity because secondary reinforcement is obstructed.

Compare learning across longer delays: Learning should decrease less in the supported condition than in the obstructed condition, because the supported condition includes a more immediate reinforcing influence during the interval.

The condition favoring secondary reinforcement should show the smaller decrease in learning as delay increases.

Key Takeaways

  1. A delayed primary reward does not have to be the only influence operating during a delay.
  2. Regularly occurring stimuli during the delay can favor the development of secondary reinforcement.
  3. Secondary reinforcement provides a more immediate reinforcing influence and can extend a goal gradient across time.
  4. Hull's proposal connects secondary reinforcement with molar stimulus traces in explaining how a goal gradient can span time.
  5. Learning decreases with increased delay less when conditions favor secondary reinforcement than when those conditions obstruct it.

Key Takeaways

  • Secondary reinforcement can operate during the interval between an action and a delayed primary reward.
  • Regularly occurring stimuli can support this secondary reinforcement and provide a more immediate influence during the delay.
  • Together with molar stimulus traces, secondary reinforcement can extend a goal gradient beyond the period that traces alone would cover.
  • The effect of delay depends partly on whether conditions support or obstruct secondary reinforcement.
  • The key comparison is between delays with different opportunities for secondary reinforcement, not merely between delay and no delay.