Temporal-Difference Error
Eligibility traces preserve decayed information about earlier state visits and actions during an episode.
Learning from the Current Step
Temporal-difference learning is a class of methods for solving finite Markov decision problems. Its defining idea is that learning can make progress step by step rather than waiting for a complete outcome. It also does not require a model of the environment. In an actor-critic method with eligibility traces, each observed step produces a temporal-difference error, updates two traces, and uses those traces to change two corresponding sets of weights.
The TD error is the signal that connects an observed transition to updates of both the policy weights and the state-value weights.
One Error, Two Update Paths
An actor-critic method contains two parameterized parts. The policy parameterization represents the policy, while the state-value parameterization represents estimated state values. Eligibility traces extend this arrangement by maintaining one trace for each set of weights: a policy trace for the policy weights and a state-value trace for the state-value weights.
w ← w + βδew
θ ← θ + αδeθ
Traces Through Earlier Visits
Eligibility traces preserve decayed information about earlier state visits and actions during an episode. A trace therefore carries some memory of what happened before the current step. The information from earlier visits is not retained at full strength indefinitely: it decays as the episode progresses. In an actor-critic method, the state-value trace and policy trace preserve this history for their respective weight updates.
Following One Earlier Visit
Consider an episode in which a state and an action occur, and later another step produces a TD error.
Earlier visit: The policy trace and state-value trace record information associated with the earlier state visit and action.
Progress to a later step: As the episode continues, the earlier information remains available through the traces but is decayed.
Weight update: When the later TD error is used, the traces carry information from the earlier visit into the policy-weight and state-value-weight updates.
Eligibility traces let a later TD error affect updates using information from earlier state visits and actions.
Policy and Value Parameterizations
| Part of the method | What it represents | Trace | Weight update |
|---|---|---|---|
| Policy parameterization | The policy | Policy trace using the gradient of the log policy | θ ← θ + αδeθ |
| State-value parameterization | Estimated state values | State-value trace using the gradient of the estimated state-value function | w ← w + βδew |
The policy and state-value parts should not be treated as interchangeable. The state-value trace uses the gradient of the estimated state-value function. The policy trace uses the gradient of the log policy. These different gradients reflect the different parameterizations. Both updates use the same TD error, but each update applies that error through the trace belonging to its own parameterization.
Model-Free Incremental Learning
A method that requires no model can learn without being supplied a complete and accurate description of the environment. Temporal-difference learning is described as not requiring a model of the environment.
Fully incremental computation means that learning can proceed step by step. Temporal-difference learning can make progress from ongoing experience instead of waiting for a complete outcome before making progress.
Choosing Among Three Method Classes
| Method class | Model requirement | Incremental computation | Main strength | Main weakness or trade-off |
|---|---|---|---|---|
| Dynamic programming | Requires a complete and accurate model | Not identified here as the defining feature | Mathematically well developed | Depends on having the required model |
| Monte Carlo methods | Does not require a model | Not well suited to step-by-step incremental computation | Conceptually simple | Less suited to incremental computation |
| Temporal-difference learning | Does not require a model | Fully incremental | Can learn step by step without a model | More complex to analyze |
The comparison should not be reduced to a claim that one method is always fastest or best. The source notes that the methods differ in efficiency and speed of convergence but does not give a universal ranking. The clearest distinction for temporal-difference learning is its combination of no model requirement and fully incremental computation, balanced against greater analytical complexity.
Auditing an Update
A learner reports that a later TD error changed both the policy weights and the state-value weights. Explain the four-part audit you would perform before deciding whether the update is correct.
Hints
- Begin with the observed transition.
- Then inspect the TD error and the two trace updates.
- Finish by checking each weight update against its corresponding trace and parameter set.
A Four-Part Debugging Order
Determine the reliable order for inspecting an actor-critic update with eligibility traces.
1. Transition: Inspect the observed transition that supplied the current step's information.
2. TD error: Inspect the temporal-difference error computed from that step.
3. Trace updates: Inspect the state-value and policy eligibility traces, including the decayed information they carry from earlier visits and actions.
4. Weight updates: Inspect the state-value update using w ← w + βδew and the policy update using θ ← θ + αδeθ.
The reliable debugging order is transition, TD error, trace updates, and weight updates.
Common Update Mistakes
Treating the policy trace and state-value trace as one shared trace.
The actor-critic method maintains one trace for the policy weights and another for the state-value weights.
Fix:
Track the policy trace and state-value trace separately.Using the same gradient description for both traces.
The state-value trace uses that gradient, while the policy trace uses the gradient of the log policy.
Fix:
Match each trace to its own parameterization and gradient.Applying the state-value update to the policy weights, or the policy update to the state-value weights.
The two equations update different weight sets.
Fix:
Use w ← w + βδew for state-value weights and θ ← θ + αδeθ for policy weights.Assuming that no model means no observed experience is needed.
The defining claim is that TD learning does not require a model; its process still proceeds from observed steps.
Fix:
Distinguish learning from experience from learning through a supplied model.Assuming that a no-model method is automatically fully incremental.
Monte Carlo methods do not require a model but are not well suited to step-by-step incremental computation.
Fix:
Check model requirement and incremental computation as separate properties.
When debugging, inspect the transition, TD error, trace updates, and weight updates in that order. This order follows the flow of information through the method and helps separate an incorrect learning signal from an incorrect trace or weight update.
Key Takeaways
- Eligibility traces preserve decayed information about earlier state visits and actions during an episode.
- An actor-critic method maintains separate policy and state-value traces.
- The state-value trace uses the gradient of the estimated state-value function, while the policy trace uses the gradient of the log policy.
- The TD error scales updates to both parameter sets: w ← w + βδew and θ ← θ + αδeθ.
- Temporal-difference learning requires no model and is fully incremental, while dynamic programming requires a complete and accurate model and Monte Carlo methods are not well suited to step-by-step incremental computation.
Key Takeaways
- Eligibility traces carry decayed information from earlier state visits and actions into later updates.
- The policy and state-value parameterizations have different roles, traces, gradients, and weight updates.
- The TD error is the shared signal that scales both policy-weight and state-value-weight changes.
- Temporal-difference learning is model-free and fully incremental.
- Dynamic programming, Monte Carlo methods, and TD learning trade off model requirements, incremental computation, conceptual simplicity, and analytical complexity.