Concepts / Value Functions and Returns

Value Functions and Returns

The task is simplified by assuming that action selection has already been learned.

  • Programming

Why the Task Is Simplified

An agent normally faces two related learning problems: deciding what to do and learning what outcomes to expect. To make the relationship between TD error and dopamine neuron activity easier to examine, this task temporarily removes the first problem. Action selection is assumed to have already been learned, so the remaining task is to predict the future return associated with the states the agent encounters.

assumed learnedremainsAction selectionFixed policyFuture-returnpredictionFuture-returnprediction
What parts of the full learning problem are removed when action selection is already learned?

Prediction Under a Fixed Policy

A policy specifies how the agent acts. In this setup, the policy is fixed: it already supplies the actions that obtain reward. Because those action choices do not change while value learning occurs, the learner evaluates the consequences of a known pattern of behavior. The value function therefore predicts the future return from states encountered while that fixed policy is followed.

This is called a prediction task or policy evaluation. The prediction concerns what future reward is likely to follow a state, while the policy determines which actions produce the subsequent experience. Policy evaluation is therefore different from learning which action should be selected. It evaluates the outcomes of actions that the already learned policy chooses.

encountered bydetermines actionis predicted asCurrent stateencountered situationFixed policyaction choiceFuture outcomereward consequencesValue predictionfuture return
How does choosing actions according to a fixed policy determine which future rewards each state or cue is predicted to produce?

A fixed policy changes what is being predicted

Consider a generated deterministic task with two encountered situations, State A and State B. The policy already chooses an action for each situation, and those choices are not being learned during this analysis.

Identify the behavior: The fixed policy supplies the action taken after State A and after State B.

Follow the consequences: Because the task is deterministic, following the fixed policy determines the future reward consequences associated with each state.

Learn the predictions: TD(0) learns a value for State A and a value for State B. These values concern the future returns produced when the fixed policy is followed.

The learner is evaluating the fixed policy, not discovering a new action-selection rule.

learnsusesAction learningchoose what to doChanging policyaction choice is learnedPolicy evaluationpredict future returnFixed policyaction choice is supplied
What is the difference between learning which action to take and learning the reward prediction associated with actions chosen by an already learned policy?

Representing States and Values

The simplified setup also makes a specific choice about representation. It uses a CSC representation in which each state visited at each time step in a trial has its own separate internal stimulus. In effect, each relevant state-time situation is treated as a separate item. This makes the representation equivalent to the tabular case.

The value function is implemented with TD(0) in a lookup table. The table contains a value prediction for each represented state-time entry, and it starts with zero for every state. When a state is encountered, its separate CSC input identifies which table entry is relevant for the prediction.

identifies entryidentifies entryCSC State Aseparate stimulusState Astored valueCSC State Bseparate stimulusState Bstored value
How does a CSC input identify the current situation and map to a value stored in the TD(0) lookup table?

Returns Across Time

A value prediction concerns the future return following the current state when the fixed policy is followed. The return is therefore connected to what happens over the subsequent sequence of states and rewards, not merely to the current state in isolation. TD(0) learns these predictions from the states generated by the fixed policy.

For this discussion, the task is deterministic and the discount factor is very nearly one, so discounting can be ignored. These assumptions remove additional sources of complexity and keep attention on the central question: how the value associated with a represented state relates to the future reward expected under the fixed policy.

followed byproduces consequencescontributes toCurrent statevalue predictionNext statefixed-policy experienceRewardlater outcomeFuture returnpredicted from currentstate
How do rewards received over time contribute to the future return predicted for the current situation?

The value stored for a state is not a prediction of an action choice in this simplified setup. It is a prediction of the future return that follows that state when the fixed policy determines what happens next.

Mistakes to Avoid

  • Treating policy evaluation as action learning.

    The simplified task assumes that action selection has already been learned and focuses only on predicting future returns under the fixed policy.

    Fix: Ask which actions the fixed policy supplies, then ask what future return those actions produce.

  • Forgetting that the policy is fixed.

    The value function is evaluating the consequences of the policy being followed, rather than changing that policy during the analysis.

    Fix: Keep the policy constant while interpreting the learned values.

  • Confusing the CSC representation with the learning algorithm.

    CSC specifies how state-time situations are represented, whereas TD(0) is the learning algorithm used with the lookup table.

    Fix: Describe CSC as the input representation and TD(0) as the value-learning method.

  • Assuming that the lookup table starts with learned values.

    The simplified setup specifies that the TD(0) lookup table is initialized with zero for every state.

    Fix: Distinguish initial stored values from the predictions learned as experience is processed.

Check Your Understanding

MEDIUM

Explain, in your own words, why fixing the policy makes the task a prediction problem. Then identify the separate roles of the CSC representation and the TD(0) lookup table.

Hints
  • Start by stating what the fixed policy supplies.
  • Then describe what the value function predicts.
  • Finally, separate the way a state is represented from the method that learns its stored value.

What do you think happens?

If action selection has already been learned, what remains for TD(0) to learn?

  • Which action to select next
  • The future-return value of encountered states under the fixed policy
  • A new CSC representation for every state
Reveal answer

Answer: The future-return value of encountered states under the fixed policy

The simplified setup removes the action-selection problem. TD(0) learns value predictions stored in the lookup table for states encountered while the fixed policy is followed.

The Essential Picture

  1. Action selection is assumed to be learned so that the analysis can focus on prediction.
  2. A fixed policy determines the actions and consequences whose future returns are being predicted.
  3. Policy evaluation is a prediction task, not a simultaneous action-selection problem.
  4. CSC gives each relevant state-time situation a separate input representation.
  5. TD(0) learns value predictions in a lookup table that starts with zero for every state.

Key Takeaways

  • The task is simplified by treating action selection as already learned.
  • The remaining problem is to predict future returns from states encountered under a fixed policy.
  • The fixed policy determines the behavior whose consequences the value function evaluates.
  • CSC separates relevant state-time situations, while TD(0) learns their values in a zero-initialized lookup table.