Model-Free and Model-Based Reinforcement Learning
Terminology must be interpreted within the field that uses it.
Same Words, Different Questions
A learner can understand the everyday meanings of positive, negative, action, and control and still misunderstand a reinforcement learning text. The problem is not that one field uses the words correctly and another uses them incorrectly. Psychology and reinforcement learning use the same words to organize different questions. Psychology often describes how behavior changes under stimuli and consequences. Reinforcement learning uses a more abstract vocabulary for how an agent selects actions and influences its environment.
| Term | Psychological use described in the source | Reinforcement-learning use |
|---|---|---|
| Negative reinforcement | Removing an aversive stimulus while behavior increases | A signal with a negative numerical value; the sign alone does not establish that an aversive stimulus was removed |
| Action | May be distinguished from related ideas such as decision or response | A broad category for what the agent selects and uses to influence the environment |
| Control | Usually emphasizes stimulus control | Emphasizes the agent's influence on the environment |
Signals Are Not Behavioral Definitions
In reinforcement learning, reward and reinforcement signals can have positive or negative numerical values. The sign describes the value assigned to the signal. It does not, by itself, identify the psychological process that produced a change in behavior. In particular, a negative reinforcement-learning value does not by itself mean that an aversive stimulus was removed, and it does not establish the behavioral definition of negative reinforcement.
The word action is also broader in reinforcement learning than in some psychological accounts. Reinforcement learning uses action as a broad category rather than separating every distinction among action, decision, and response. Similarly, control emphasizes the agent's influence on the environment, whereas psychology usually emphasizes stimulus control.
A single positive-or-negative reward scale is useful as an abstraction, but it can hide an important distinction. Appetitive and aversive systems have qualitatively different properties and involve different brain mechanisms. An action can therefore involve both appetitive and aversive considerations, while one signed number compresses those systems into a single scale.
Cached Choices Without a World Model
A model-free decision uses cached values or a cached policy rather than an environment model. The agent selects an action by reading what has already been stored for the current state; it does not use a model to simulate possible future routes during that decision.
A Cached Maze Choice
At state S1 in the maze task, distinguish the model-free justification for choosing left from the action that is chosen.
Start at S1: The agent is at the current state and must select an action.
Read cached guidance: A model-free strategy retrieves an experienced action choice through a cached value or cached policy.
Choose left: The selected action is L. The model-free explanation is the stored guidance, not a newly simulated route.
The model-free strategy chooses L at S1 by retrieving cached guidance rather than evaluating possible future sequences with an environment model.
The Two Parts of an Environment Model
Model-based reinforcement learning uses an environment model with two connected parts. The state-transition model represents what state follows from a state-action choice. In the maze description, this component can be represented as a decision tree. The reward model associates distinctive goal-box features with the rewards available in those boxes. Rewards associated with S1, S2, and S3 are also part of the reward model, but they are zero in the described task and are not shown.
| Model component | Question it answers | Maze role |
|---|---|---|
| State-transition model | What state follows from this state-action choice? | Represents the states reached after choosing left or right |
| Reward model | What reward is associated with the resulting goal-box features? | Associates goal-box features with available rewards |
The transition and reward components answer different questions but work together during planning.
Planning by Simulated Routes
A model-based agent uses its environment model to simulate several action choices. It follows the predicted consequences of one choice, continues through the possible sequence, identifies the terminal outcome, and compares that path's predicted return with the returns of other simulated paths. This comparison is planning. The selected action is the one that belongs to the sequence with the highest predicted return.
Comparing Paths in the Maze
A rat begins at S1. It can choose left or right, reach S2 or S3, and then choose left or right again. Terminal goal boxes provide associated rewards. How does a model-based strategy reach its choice?
Represent the first choice: The agent uses the transition model to predict which state follows a left or right choice from S1.
Extend each route: For each predicted state, the agent simulates the next left or right choice and follows the possible sequence to a terminal goal box.
Attach predicted rewards: The reward model identifies the reward associated with the terminal goal-box features.
Compare routes: The agent compares the predicted returns of the simulated paths.
Select the first action: The first action belonging to the sequence with the highest predicted return is selected.
Model-based selection is a plan over predicted transitions and rewards, not merely a retrieved response to S1.
One Maze, Two Justifications
| Question | Model-free decision | Model-based decision |
|---|---|---|
| What is used at S1? | A cached value or cached policy | An environment model |
| What happens before choosing? | Stored guidance is retrieved | Possible routes are simulated |
| What is selected in the source task? | L at S1, then R at S2 | L at S1, then R at S2 |
| What distinguishes the strategy? | The decision reads cached guidance | The decision evaluates possible sequences |
Check Your Understanding
A navigation agent is at S1. It chooses L because a stored policy associates S1 with L. It does not consult predicted transitions or predicted rewards during this choice. Is this decision model-free or model-based, and why?
Hints
- Ask whether the agent uses an environment model.
- Look for cached guidance versus simulated routes.
A second agent at S1 predicts the states reached by L and R, follows each possible route to a terminal goal box, uses the reward model to identify the associated rewards, and compares the predicted returns. What makes this decision model-based?
Hints
- Identify the transition component.
- Identify the reward component.
- Find the comparison of simulated paths.
Treating a negative reinforcement-learning signal as psychological negative reinforcement
The reinforcement-learning sign describes numerical value. Psychological negative reinforcement concerns removing an aversive stimulus while behavior increases.
Fix:
Interpret the signal within reinforcement learning and do not infer the psychological process from its sign alone.Assuming that different routes must reveal different strategies
Both model-free and model-based approaches select L at S1 in the source maze task.
Fix:
Ask whether the choice came from cached guidance or from simulated transitions and rewards.Describing model-free selection as planning
Model-free decisions use cached values or a cached policy rather than an environment model.
Fix:
Describe the stored value or policy that guides the current action.Reducing appetitive and aversive systems to a complete single-scale description
The source notes that appetitive and aversive systems have qualitatively different properties and involve different brain mechanisms.
Fix:
Treat a signed reward as a useful abstraction that may omit the distinction between systems.
Decision Process Is the Key
- Psychological negative reinforcement is a behavioral process involving removal of an aversive stimulus and increased behavior; a negative reinforcement-learning signal is a numerical value.
- Reinforcement learning uses action broadly and emphasizes the agent's influence on the environment, unlike psychology's usual emphasis on stimulus control.
- Model-free decisions select actions from cached values or a cached policy without using an environment model.
- A model-based environment model contains a state-transition model and a reward model.
- Model-based planning simulates possible routes, predicts their outcomes and rewards, compares their predicted returns, and selects the action belonging to the best sequence.
Key Takeaways
- Terminology must be interpreted within the field using it.
- A reinforcement-learning signal's numerical sign is not the same thing as the psychological definition of positive or negative reinforcement.
- Model-free behavior retrieves cached guidance, while model-based behavior evaluates simulated action sequences.
- The transition model predicts where choices lead, and the reward model identifies rewards associated with resulting outcomes.
- The same navigation route can be selected by both strategies; their decision processes distinguish them.