Value Functions and Returns
The task is simplified by assuming that action selection has already been learned.
Why the Task Is Simplified
An agent normally faces two related learning problems: deciding what to do and learning what outcomes to expect. To make the relationship between TD error and dopamine neuron activity easier to examine, this task temporarily removes the first problem. Action selection is assumed to have already been learned, so the remaining task is to predict the future return associated with the states the agent encounters.
Prediction Under a Fixed Policy
A policy specifies how the agent acts. In this setup, the policy is fixed: it already supplies the actions that obtain reward. Because those action choices do not change while value learning occurs, the learner evaluates the consequences of a known pattern of behavior. The value function therefore predicts the future return from states encountered while that fixed policy is followed.
This is called a prediction task or policy evaluation. The prediction concerns what future reward is likely to follow a state, while the policy determines which actions produce the subsequent experience. Policy evaluation is therefore different from learning which action should be selected. It evaluates the outcomes of actions that the already learned policy chooses.
A fixed policy changes what is being predicted
Consider a generated deterministic task with two encountered situations, State A and State B. The policy already chooses an action for each situation, and those choices are not being learned during this analysis.
Identify the behavior: The fixed policy supplies the action taken after State A and after State B.
Follow the consequences: Because the task is deterministic, following the fixed policy determines the future reward consequences associated with each state.
Learn the predictions: TD(0) learns a value for State A and a value for State B. These values concern the future returns produced when the fixed policy is followed.
The learner is evaluating the fixed policy, not discovering a new action-selection rule.
Representing States and Values
The simplified setup also makes a specific choice about representation. It uses a CSC representation in which each state visited at each time step in a trial has its own separate internal stimulus. In effect, each relevant state-time situation is treated as a separate item. This makes the representation equivalent to the tabular case.
The value function is implemented with TD(0) in a lookup table. The table contains a value prediction for each represented state-time entry, and it starts with zero for every state. When a state is encountered, its separate CSC input identifies which table entry is relevant for the prediction.
Returns Across Time
A value prediction concerns the future return following the current state when the fixed policy is followed. The return is therefore connected to what happens over the subsequent sequence of states and rewards, not merely to the current state in isolation. TD(0) learns these predictions from the states generated by the fixed policy.
For this discussion, the task is deterministic and the discount factor is very nearly one, so discounting can be ignored. These assumptions remove additional sources of complexity and keep attention on the central question: how the value associated with a represented state relates to the future reward expected under the fixed policy.
The value stored for a state is not a prediction of an action choice in this simplified setup. It is a prediction of the future return that follows that state when the fixed policy determines what happens next.
Mistakes to Avoid
Treating policy evaluation as action learning.
The simplified task assumes that action selection has already been learned and focuses only on predicting future returns under the fixed policy.
Fix:
Ask which actions the fixed policy supplies, then ask what future return those actions produce.Forgetting that the policy is fixed.
The value function is evaluating the consequences of the policy being followed, rather than changing that policy during the analysis.
Fix:
Keep the policy constant while interpreting the learned values.Confusing the CSC representation with the learning algorithm.
CSC specifies how state-time situations are represented, whereas TD(0) is the learning algorithm used with the lookup table.
Fix:
Describe CSC as the input representation and TD(0) as the value-learning method.Assuming that the lookup table starts with learned values.
The simplified setup specifies that the TD(0) lookup table is initialized with zero for every state.
Fix:
Distinguish initial stored values from the predictions learned as experience is processed.
Check Your Understanding
Explain, in your own words, why fixing the policy makes the task a prediction problem. Then identify the separate roles of the CSC representation and the TD(0) lookup table.
Hints
- Start by stating what the fixed policy supplies.
- Then describe what the value function predicts.
- Finally, separate the way a state is represented from the method that learns its stored value.
What do you think happens?
If action selection has already been learned, what remains for TD(0) to learn?
Reveal answer
Answer: The future-return value of encountered states under the fixed policy
The simplified setup removes the action-selection problem. TD(0) learns value predictions stored in the lookup table for states encountered while the fixed policy is followed.
The Essential Picture
- Action selection is assumed to be learned so that the analysis can focus on prediction.
- A fixed policy determines the actions and consequences whose future returns are being predicted.
- Policy evaluation is a prediction task, not a simultaneous action-selection problem.
- CSC gives each relevant state-time situation a separate input representation.
- TD(0) learns value predictions in a lookup table that starts with zero for every state.
Key Takeaways
- The task is simplified by treating action selection as already learned.
- The remaining problem is to predict future returns from states encountered under a fixed policy.
- The fixed policy determines the behavior whose consequences the value function evaluates.
- CSC separates relevant state-time situations, while TD(0) learns their values in a zero-initialized lookup table.