Action Values and Policies in Reinforcement Learning
Model-free decisions use cached values or a cached policy rather than an environment model.
The Same Route, Different Reasoning
Imagine a rat beginning at state S1 in a maze. From S1, it can choose left or right, reach either S2 or S3, and then choose left or right again. Each terminal goal box delivers an associated reward. In the source task, both a model-free strategy and a model-based strategy select L at S1 and then R at S2. The important distinction is not the route itself. It is how the agent justifies the next action.
A model-free decision retrieves cached information. A model-based decision uses an environment model to evaluate possible sequences before acting.
Cached Action Values
In a model-free decision, the agent does not use an environment model to work out what will happen next. Instead, it uses cached values associated with actions in the current state. The current state is S1 in the maze example. The agent consults the cached action information for S1 and selects an action from that information. In the source task, this process leads to selecting L at S1 and later R at S2.
The key operation is retrieval: the agent uses cached action information for the current state rather than simulating a route through an environment model.
Cached Policies
A model-free agent can also use a cached policy. In this description, the policy provides the action to take for the current state. The policy therefore maps the current state directly to an action choice. At S1 in the maze task, the cached policy gives L; after reaching S2, it gives R.
The Environment Model
Model-based reinforcement learning uses an environment model with two connected parts. The state-transition model represents what state follows from a state-action choice. In the maze description, this part is represented as a decision tree. The reward model associates distinctive goal-box features with the rewards available in those boxes. Rewards associated with S1, S2, and S3 are also part of the reward model, but they are zero in the described task and are not shown.
| Model component | Question it answers | Maze role |
|---|---|---|
| State-transition model | What state follows from this state-action choice? | Represents the possible maze routes as a decision tree. |
| Reward model | What reward is associated with the resulting goal-box features? | Associates goal-box features with their available rewards. |
The two model components provide different information, and together they support route evaluation.
Simulating Candidate Routes
A model-based agent can compare several possible action sequences before choosing. Starting at S1, it considers a first action such as L or R. The transition model predicts the state reached by that choice. The agent then continues through the possible sequence, using the transition model to follow predicted consequences. When the sequence reaches a terminal goal box, the reward model identifies the reward associated with that outcome. The agent compares the predicted return of this sequence with the predicted returns of other simulated sequences.
The maze choice at S1
Trace how the model-based strategy selects the first action in the source maze task.
Represent alternatives: At S1, the agent considers the available left and right choices as possible beginnings of routes.
Predict consequences: The state-transition model predicts the states reached by the choices and continues the possible sequences through the maze.
Evaluate terminal outcomes: When a simulated sequence reaches a terminal goal box, the reward model identifies the reward associated with that outcome.
Compare routes: The agent compares the predicted return of each simulated route. This comparison is a simple form of planning.
Select the first action: In the source task, the sequence with the highest predicted return begins with L, so the model-based agent selects L at S1.
The model-based agent selects L at S1 by comparing predicted route returns, not by merely retrieving an action associated with S1.
Planning is the comparison of predicted returns from simulated routes. The selected action is the action that belongs to the sequence with the highest predicted return.
Reading the Difference
| Question | Model-free decision | Model-based decision |
|---|---|---|
| What information is consulted? | Cached action values or a cached policy. | A state-transition model and a reward model. |
| How is the action selected? | Retrieve the action associated with the current state. | Simulate possible sequences and compare their predicted returns. |
| What happens at S1 in the source task? | Select L from cached information. | Select L after evaluating simulated routes. |
| What happens at S2 in the source task? | Select R from cached information. | Select R as part of the route with the highest predicted return. |
Common Mistakes
Assuming that a route proves an agent is model-based.
The same route can be produced by cached information or by simulated route comparison.
Fix:
Inspect the decision process: retrieval indicates model-free behavior, while predicted state transitions, rewards, and returns indicate model-based behavior.Treating the transition model and reward model as the same component.
The source distinguishes the questions answered by the two components.
Fix:
Use the transition model to predict where a choice leads and the reward model to associate the resulting goal-box features with rewards.Calling cached action selection planning.
Model-free behavior retrieves a cached value or policy rather than evaluating possible sequences with an environment model.
Fix:
Reserve planning for the comparison of predicted returns from simulated routes.Looking only at the first action and ignoring how it was chosen.
The action is identical, but the justifications can differ.
Fix:
Ask whether the agent retrieved L from cached information or selected it after comparing simulated consequences.
Check Your Understanding
An agent is at S1 in the maze. It consults information already cached for S1 and selects L without using a state-transition model or a reward model. Is this decision model-free or model-based? What evidence supports your answer?
Hints
- Identify whether the agent retrieved cached information or simulated possible routes.
- A model-free decision uses cached action values or a cached policy.
A second agent at S1 uses a decision-tree-like state-transition model to follow possible routes, uses a reward model for terminal goal-box outcomes, and compares predicted returns. What makes this planning, and which first action does the source task say it selects?
Hints
- Planning is the comparison of predicted returns from simulated routes.
- The source task says the selected route begins with L.
Key Takeaways
- Model-free reinforcement learning selects actions from cached action values or a cached policy instead of using an environment model.
- A model-based environment model has a state-transition component and a reward component.
- The transition component predicts the next state after a state-action choice, while the reward component associates outcomes with rewards.
- Model-based planning compares predicted returns from simulated action sequences.
- In the source maze, both approaches choose L at S1 and R at S2, but they arrive at those choices through different decision processes.
Key Takeaways
- Model-free decisions retrieve cached values or a cached policy.
- Model-based decisions use transition and reward models to simulate possible routes.
- Planning means comparing the predicted returns of those simulated routes.
- The same action sequence can be selected by both strategies, so the decision process matters more than the route alone.