Concepts / Prediction and Control in Reinforcement Learning

Prediction and Control in Reinforcement Learning

Instrumental conditioning is learning in which consequences modify behavior.

  • Programming

A Consequence Changes What Happens Next

Consider an animal that performs an action and then receives something it likes. If that action becomes more likely later, the important learning event is not simply that the animal experienced a reward. The reward followed a particular behavior, and the consequence changed the future tendency to produce that behavior.

Instrumental conditioning is learning in which consequences modify behavior.

This is a control relationship. A behavior leads to a consequence, and that consequence is associated with a change in future behavior. In reinforcement learning, the algorithmic counterpart is policy improvement: the agent's future action choices are changed in light of what follows its actions.

Contingency Makes the Difference

For the instrumental arrangement described here, a reinforcing stimulus must be contingent on the behavior. Contingent means that the consequence occurs because it follows a particular behavior in the learning arrangement, rather than being merely an event the learner experiences independently of that behavior.

leads tomodifiesdoes not establishBehavioraction occursStimulusoccurs independentlyReinforcing stimulusfollows behaviorNo behavior linknot tied to an actionFuture behaviortendency changes
What changes when a reinforcing stimulus follows a behavior rather than occurring independently?
QuestionInstrumental arrangementBroader prediction setting
What is central?A behavior and its consequenceHow an environmental feature is expected to unfold
Must the stimulus evaluate an earlier action?The reinforcing consequence is contingent on behaviorNo; a predicted stimulus need not be a reward or penalty evaluating an earlier action
What can change?Future behaviorAn estimate of an upcoming quantity

This contingency-based distinction also limits the connection with classical conditioning. Prediction algorithms and classical conditioning can both involve predicting upcoming stimuli. However, the predicted stimulus does not have to be a reward or penalty that evaluates an earlier action. Therefore, prediction is not automatically an instrumental or classical-conditioning process; its meaning depends on what is being predicted and how that quantity relates to behavior and consequences.

From Consequences to Policy Improvement

A reinforcement-learning policy describes the agent's action choices. Policy improvement changes those choices so that some actions become more likely and others become less likely. Rewards increase the tendency toward rewarded behavior, while penalties decrease the tendency toward penalized behavior.

is followed bymodifiescorresponds toActionagent behavesConsequencereward or penaltyBehavior tendencyincreases or decreasesPolicyfuture action choices
How do observed behavior-consequence outcomes cause a reinforcement-learning policy to choose some actions more often?

The diagram should be read as a control loop, not as a claim about one particular internal mechanism. The experimental consequence supplies the direction of change: a rewarded behavior becomes more likely, whereas a penalized behavior becomes less likely. Reinforcement learning's policy-improvement algorithms provide an algorithmic counterpart to this form of control.

Choosing Between Two Behaviors

An animal can perform behavior A or behavior B. When behavior A is followed by something the animal likes, behavior A becomes more likely later. When behavior B is followed by a penalty, behavior B becomes less likely later. How does this illustrate the connection between instrumental conditioning and policy improvement?

Identify the behavior: The relevant learning event is not just the receipt of a stimulus. Behavior A and behavior B are the actions whose consequences are being considered.

Identify the contingency: The liked consequence follows behavior A, and the penalty follows behavior B. Each consequence is connected to a particular behavior.

Track the behavioral change: The tendency to produce behavior A increases, while the tendency to produce behavior B decreases.

Translate to reinforcement learning: An algorithmic policy-improvement process changes future action choices in the same directional pattern: it favors the rewarded behavior and disfavors the penalized behavior.

Instrumental conditioning describes the behavior-consequence control relationship; policy improvement is its reinforcement-learning algorithmic counterpart.

Prediction Looks Toward the Future

A prediction algorithm in reinforcement learning estimates a quantity whose value depends on how features of the environment are expected to unfold in the future.

A useful way to understand prediction is to follow one target quantity through time. First identify the environmental feature to be predicted. Next consider how that feature is expected to unfold as the agent continues interacting with the environment. Finally produce an estimate of the future-dependent quantity.

guidesproduces expectedis estimated asCurrent policyactions selectedFuture interactionenvironment unfoldsFuture rewardtarget quantityPolicy evaluationestimated expected reward
How does a prediction of future rewards estimate how good the current policy is?

When a prediction algorithm estimates future reward under a particular policy, it is performing policy evaluation. The estimate answers a question about the policy being followed: what amount of future reward is expected as the agent continues interacting with the environment?

Reward Is One Prediction Target

Prediction targetWhat is estimated?Relation to policy evaluation
Future rewardThe reward expected as the agent continues interacting with the environmentWhen estimated under a policy, this is policy evaluation
Upcoming stimulusHow a stimulus is expected to unfold in the futureIt need not evaluate an earlier action
Other numerical-valued environmental featureThe expected future behavior of that featureIt is prediction, but not necessarily future-reward evaluation

Future reward is the main reinforcement-learning example, but it is not the boundary of prediction. Prediction algorithms can estimate upcoming stimuli and other numerical-valued environmental features. The important question is not whether the target is called a reward; it is whether the target is a quantity whose expected future behavior can be estimated.

estimated under policyboth are possible targetsFuture rewardnumerical targetUpcoming stimulusenvironmental featurePolicy evaluationunder a policyOther featurenumerical value
What is being predicted in each case, and when does a prediction count as policy evaluation?

Why Prediction Supports Control

Many reinforcement-learning questions concern what has not happened yet. An agent may need an estimate of the reward it can expect while it continues interacting with the environment. That estimate provides information about the consequences associated with the policy being evaluated.

is estimated bysupportscontributes toFuture rewardtarget quantityPredictionexpected future valuePolicy evaluationcurrent policy assessedPolicy improvementfuture choices change
How do predicted future outcomes provide information that an agent can use to compare actions and improve its policy?

The connection is indirect but important. A prediction algorithm estimates future reward associated with the policy being evaluated. Policy evaluation is an integral component of algorithms for improving policies, so prediction can support improvement without being identical to improvement itself.

  • Treating reward experience as sufficient for instrumental conditioning

    The relevant relationship is that the reward followed a particular behavior and that the behavior later changed.

    Fix: Identify both the behavior and the consequence, then ask whether the consequence was contingent on that behavior.

  • Using prediction and policy improvement as synonyms

    Future-reward estimation is policy evaluation; evaluation contributes to, but is not identical to, policy improvement.

    Fix: Separate estimating what a policy produces from changing the policy's future action choices.

  • Assuming every prediction is a reward prediction

    Prediction algorithms can estimate upcoming stimuli and other numerical-valued environmental features.

    Fix: Name the target quantity first, then determine whether it is future reward, an upcoming stimulus, or another numerical-valued feature.

  • Assuming every prediction task is a classical-conditioning task

    The connection with classical conditioning is narrower: both can involve predicting upcoming stimuli, but a predicted stimulus need not evaluate an earlier action.

    Fix: Describe the predicted quantity and its relationship to behavior before assigning a conditioning interpretation.

Check the Mechanism

MEDIUM

A learner observes that an agent repeatedly performs behavior X. Something the agent likes follows behavior X, and behavior X becomes more likely later. Explain the case using all four terms: behavior, contingency, prediction, and policy improvement.

Hints
  • Start by identifying what happens before the future tendency changes.
  • Ask whether the liked consequence follows a particular behavior.
  • Decide whether the case describes an estimate of future reward, a change in action tendency, or both.
  1. Instrumental conditioning is learning in which consequences modify behavior. In the arrangement described here, the reinforcing consequence is contingent on a particular behavior. Reinforcement learning expresses a corresponding form of control through policy improvement, in which future action tendencies change. A prediction algorithm estimates how a target quantity is expected to unfold in the future. When the target is future reward under a policy, the prediction is policy evaluation. Evaluation supports policy improvement but is not the same operation. Prediction is broader than reward estimation because its targets can also include upcoming stimuli and other numerical-valued environmental features.

Key Takeaways

  • Instrumental conditioning links a behavior to a consequence that modifies future behavior.
  • Contingency is central: the reinforcing consequence follows the relevant behavior rather than occurring independently of it.
  • Policy improvement is the reinforcement-learning algorithmic counterpart to consequence-based control.
  • Prediction estimates future-dependent quantities; future reward prediction under a policy is policy evaluation.
  • Prediction is broader than reward estimation and can target upcoming stimuli or other numerical-valued environmental features.