Concepts / Model-Free and Model-Based Reinforcement Learning

Model-Free and Model-Based Reinforcement Learning

Terminology must be interpreted within the field that uses it.

  • Programming

Same Words, Different Questions

A learner can understand the everyday meanings of positive, negative, action, and control and still misunderstand a reinforcement learning text. The problem is not that one field uses the words correctly and another uses them incorrectly. Psychology and reinforcement learning use the same words to organize different questions. Psychology often describes how behavior changes under stimuli and consequences. Reinforcement learning uses a more abstract vocabulary for how an agent selects actions and influences its environment.

TermPsychological use described in the sourceReinforcement-learning use
Negative reinforcementRemoving an aversive stimulus while behavior increasesA signal with a negative numerical value; the sign alone does not establish that an aversive stimulus was removed
ActionMay be distinguished from related ideas such as decision or responseA broad category for what the agent selects and uses to influence the environment
ControlUsually emphasizes stimulus controlEmphasizes the agent's influence on the environment
describesexaminesorganizesorganizesPsychologyBehavior changeAction selectionStimuli andconsequencesReinforcementlearningEnvironment influence
How do the same words organize different questions in psychology and reinforcement learning?

Signals Are Not Behavioral Definitions

In reinforcement learning, reward and reinforcement signals can have positive or negative numerical values. The sign describes the value assigned to the signal. It does not, by itself, identify the psychological process that produced a change in behavior. In particular, a negative reinforcement-learning value does not by itself mean that an aversive stimulus was removed, and it does not establish the behavioral definition of negative reinforcement.

The word action is also broader in reinforcement learning than in some psychological accounts. Reinforcement learning uses action as a broad category rather than separating every distinction among action, decision, and response. Similarly, control emphasizes the agent's influence on the environment, whereas psychology usually emphasizes stimulus control.

A single positive-or-negative reward scale is useful as an abstraction, but it can hide an important distinction. Appetitive and aversive systems have qualitatively different properties and involve different brain mechanisms. An action can therefore involve both appetitive and aversive considerations, while one signed number compresses those systems into a single scale.

may engagemay engagecompressed intocompressed intomay omitActionAppetitive systemSigned valueSystem distinctionAversive system
Why can one signed reward number fail to preserve the distinction between appetitive and aversive systems?

Cached Choices Without a World Model

A model-free decision uses cached values or a cached policy rather than an environment model. The agent selects an action by reading what has already been stored for the current state; it does not use a model to simulate possible future routes during that decision.

look upprovide stored guidanceselectCurrent stateCached values orpolicyAction choiceSelected action
How does a model-free agent move from the current state to an action without simulating future outcomes?

A Cached Maze Choice

At state S1 in the maze task, distinguish the model-free justification for choosing left from the action that is chosen.

Start at S1: The agent is at the current state and must select an action.

Read cached guidance: A model-free strategy retrieves an experienced action choice through a cached value or cached policy.

Choose left: The selected action is L. The model-free explanation is the stored guidance, not a newly simulated route.

The model-free strategy chooses L at S1 by retrieving cached guidance rather than evaluating possible future sequences with an environment model.

The Two Parts of an Environment Model

Model-based reinforcement learning uses an environment model with two connected parts. The state-transition model represents what state follows from a state-action choice. In the maze description, this component can be represented as a decision tree. The reward model associates distinctive goal-box features with the rewards available in those boxes. Rewards associated with S1, S2, and S3 are also part of the reward model, but they are zero in the described task and are not shown.

Model componentQuestion it answersMaze role
State-transition modelWhat state follows from this state-action choice?Represents the states reached after choosing left or right
Reward modelWhat reward is associated with the resulting goal-box features?Associates goal-box features with available rewards

The transition and reward components answer different questions but work together during planning.

containspredictscontainsassociatesconnects route to rewardEnvironment modelState-transitionmodelResulting stateReward modelGoal-box reward
What does an environment model contain, and how do predicted transitions connect to predicted rewards?

Planning by Simulated Routes

A model-based agent uses its environment model to simulate several action choices. It follows the predicted consequences of one choice, continues through the possible sequence, identifies the terminal outcome, and compares that path's predicted return with the returns of other simulated paths. This comparison is planning. The selected action is the one that belongs to the sequence with the highest predicted return.

simulatepredictcompare withpredictchoose highest predicted returnAgentRoute LPredicted outcome LRoute RPredicted outcome RSelected route
How does a model-based agent simulate different action sequences, predict their outcomes and rewards, and choose between them?

Comparing Paths in the Maze

A rat begins at S1. It can choose left or right, reach S2 or S3, and then choose left or right again. Terminal goal boxes provide associated rewards. How does a model-based strategy reach its choice?

Represent the first choice: The agent uses the transition model to predict which state follows a left or right choice from S1.

Extend each route: For each predicted state, the agent simulates the next left or right choice and follows the possible sequence to a terminal goal box.

Attach predicted rewards: The reward model identifies the reward associated with the terminal goal-box features.

Compare routes: The agent compares the predicted returns of the simulated paths.

Select the first action: The first action belonging to the sequence with the highest predicted return is selected.

Model-based selection is a plan over predicted transitions and rewards, not merely a retrieved response to S1.

One Maze, Two Justifications

QuestionModel-free decisionModel-based decision
What is used at S1?A cached value or cached policyAn environment model
What happens before choosing?Stored guidance is retrievedPossible routes are simulated
What is selected in the source task?L at S1, then R at S2L at S1, then R at S2
What distinguishes the strategy?The decision reads cached guidanceThe decision evaluates possible sequences
retrieveselectconsultsimulateselect first actionS1Cached guidanceLS1Environment modelSimulated routesL
Given the same navigation choice, how does a model-free agent's cached response differ from a model-based agent's plan?

Check Your Understanding

EASY

A navigation agent is at S1. It chooses L because a stored policy associates S1 with L. It does not consult predicted transitions or predicted rewards during this choice. Is this decision model-free or model-based, and why?

Hints
  • Ask whether the agent uses an environment model.
  • Look for cached guidance versus simulated routes.
MEDIUM

A second agent at S1 predicts the states reached by L and R, follows each possible route to a terminal goal box, uses the reward model to identify the associated rewards, and compares the predicted returns. What makes this decision model-based?

Hints
  • Identify the transition component.
  • Identify the reward component.
  • Find the comparison of simulated paths.
  • Treating a negative reinforcement-learning signal as psychological negative reinforcement

    The reinforcement-learning sign describes numerical value. Psychological negative reinforcement concerns removing an aversive stimulus while behavior increases.

    Fix: Interpret the signal within reinforcement learning and do not infer the psychological process from its sign alone.

  • Assuming that different routes must reveal different strategies

    Both model-free and model-based approaches select L at S1 in the source maze task.

    Fix: Ask whether the choice came from cached guidance or from simulated transitions and rewards.

  • Describing model-free selection as planning

    Model-free decisions use cached values or a cached policy rather than an environment model.

    Fix: Describe the stored value or policy that guides the current action.

  • Reducing appetitive and aversive systems to a complete single-scale description

    The source notes that appetitive and aversive systems have qualitatively different properties and involve different brain mechanisms.

    Fix: Treat a signed reward as a useful abstraction that may omit the distinction between systems.

Decision Process Is the Key

  1. Psychological negative reinforcement is a behavioral process involving removal of an aversive stimulus and increased behavior; a negative reinforcement-learning signal is a numerical value.
  2. Reinforcement learning uses action broadly and emphasizes the agent's influence on the environment, unlike psychology's usual emphasis on stimulus control.
  3. Model-free decisions select actions from cached values or a cached policy without using an environment model.
  4. A model-based environment model contains a state-transition model and a reward model.
  5. Model-based planning simulates possible routes, predicts their outcomes and rewards, compares their predicted returns, and selects the action belonging to the best sequence.

Key Takeaways

  • Terminology must be interpreted within the field using it.
  • A reinforcement-learning signal's numerical sign is not the same thing as the psychological definition of positive or negative reinforcement.
  • Model-free behavior retrieves cached guidance, while model-based behavior evaluates simulated action sequences.
  • The transition model predicts where choices lead, and the reward model identifies rewards associated with resulting outcomes.
  • The same navigation route can be selected by both strategies; their decision processes distinguish them.