Learning to Predict and Learning to Control
Non-contingent eligibility traces do not correlate actions with consequent changes in the reward signal.
The Missing Connection
A reinforcement signal can tell a learning system that something happened to its outcome. For control, however, that signal must do more than indicate an outcome: learning must relate what an agent did to what happened to the reward afterward. Non-contingent eligibility traces fail to make this action-to-consequent-reward-change connection.
What do you think happens?
An agent performs an action, and later the common reinforcement signal changes. What connection is needed if the agent is meant to learn control?
Reveal answer
Answer: The learning process must connect the agent's earlier action with the later change in the reward signal.
Control requires actions to be evaluated in light of their consequences. Non-contingent eligibility traces do not provide that correlation, even though they can still support prediction.
Action to Consequence
The limitation becomes clearer as a sequence. First, an agent performs an action. Later, the reinforcement signal changes. A control-learning method needs to connect these two events so that the earlier action can be considered in light of the later consequence. Non-contingent eligibility traces do not correlate the action with the consequent change in the reward signal.
Tracing the required relationship
Consider a generated team-learning situation in which an agent performs an action and the common reinforcement signal changes later. What must learning establish for control?
1. Record the first event: The agent performs an action. This is the activity whose contribution may need to be evaluated.
2. Observe the later event: The common reinforcement signal changes after the action.
3. Relate the events: Control learning requires a correlation between the agent's action and the subsequent reward change.
4. Check the trace limitation: A non-contingent eligibility trace does not supply that action-to-consequent-reward-change correlation.
The trace can be adequate for prediction, but it is inadequate as sufficient support for learning control.
Critic and Actor
The actor-critic algorithm separates two jobs. The critic learns to predict the expected return. The actor learns to control, meaning that it must learn how actions should be selected. Non-contingent eligibility traces are adequate for the critic's prediction task, but they do not support the actor's control task.
| Part of the algorithm | Role | Relationship to non-contingent eligibility traces |
|---|---|---|
| Critic | Predicts the expected return | Non-contingent eligibility traces are adequate for this prediction role |
| Actor | Controls by learning how actions should be selected | Non-contingent eligibility traces do not support this control role |
Team-Level Credit
The defined team-learning setting contains a population of agents that receives a common reinforcement signal. The agents are expected to learn differentiated patterns of activity. The objective is therefore not merely for every agent to predict the same quantity. Learning must support distinct patterns of activity within the population.
Non-contingent eligibility traces are inadequate for this team setting because they do not let the learning process correlate each agent's actions with changes in the shared reward signal. Without that correlation, the common signal does not identify how individual activity should contribute to differentiated behavior.
Shared Prediction and Behavior
A population can receive a common reinforcement signal and learn a shared prediction, but that does not by itself produce differentiated agent behavior. Prediction asks what return is expected. Differentiated control requires learning how individual actions should be selected and how those actions relate to later reward changes. The shared prediction therefore does not substitute for individual action-to-consequence information.
The problem is not that a common reinforcement signal is useless. It is useful for control only when learning can relate what an agent did to what happened to the reward afterward. The problem is that non-contingent eligibility traces do not provide this individual correlation, so the shared signal cannot by itself identify differentiated contributions.
Check Your Understanding
Explain why non-contingent eligibility traces can be adequate for the critic but inadequate for the actor in a team-learning setting.
Hints
- Start by stating what the critic learns to predict.
- Then state what the actor must learn about action selection.
- Finally, identify the action-to-consequent-reward-change connection that non-contingent traces fail to provide.
A complete explanation
Give a four-part explanation of the limitation in the team setting.
Prediction role: The critic learns to predict the expected return, and non-contingent eligibility traces are adequate for this task.
Control role: The actor learns how actions should be selected, so it needs information connecting actions with their later reward consequences.
Team requirement: The team contains multiple agents receiving a common reinforcement signal and is expected to produce differentiated patterns of activity.
Failure point: Non-contingent eligibility traces do not correlate each agent's actions with changes in the shared reward signal, so they do not adequately support differentiated control.
The traces support shared prediction but not the individual action-to-consequence credit needed for team control.
Common Reasoning Errors
Treating prediction and control as the same learning problem.
The critic predicts the expected return, while the actor must learn how actions should be selected.
Fix:
Check whether the method connects an agent's action with the subsequent reward change. That connection is required for control.Assuming a common reward automatically identifies each agent's contribution.
The shared signal does not identify how each agent's activity contributed when the learning process cannot correlate individual actions with reward changes.
Fix:
Separate the existence of a common signal from the availability of individual action-to-consequence credit.Concluding that shared prediction guarantees differentiated behavior.
The team objective includes differentiated patterns of activity, not merely a shared prediction.
Fix:
Ask whether the learning process supports control of individual actions as well as prediction of the return.
Key Takeaways
- Non-contingent eligibility traces do not correlate an action with a consequent change in the reward signal.
- The critic predicts the expected return, and non-contingent eligibility traces are adequate for that prediction role.
- The actor controls by learning how actions should be selected, which requires action-to-consequence information that non-contingent traces do not provide.
- In team learning, a common reinforcement signal must support differentiated patterns of activity across agents.
- Shared prediction does not necessarily produce differentiated behavior because prediction alone does not identify how each agent's activity should contribute.
Key Takeaways
- Non-contingent eligibility traces fail to connect actions with consequent changes in the reward signal.
- They can support the critic's prediction of expected return but not the actor's control of action selection.
- A team-learning setting needs individual action-to-shared-reward correlations to support differentiated agent activity.
- A common prediction or common reinforcement signal does not by itself produce distinct behavior among agents.