Concepts / Temporal-Difference Error

Temporal-Difference Error

Eligibility traces preserve decayed information about earlier state visits and actions during an episode.

  • Programming

Learning from the Current Step

Temporal-difference learning is a class of methods for solving finite Markov decision problems. Its defining idea is that learning can make progress step by step rather than waiting for a complete outcome. It also does not require a model of the environment. In an actor-critic method with eligibility traces, each observed step produces a temporal-difference error, updates two traces, and uses those traces to change two corresponding sets of weights.

The TD error is the signal that connects an observed transition to updates of both the policy weights and the state-value weights.

One Error, Two Update Paths

An actor-critic method contains two parameterized parts. The policy parameterization represents the policy, while the state-value parameterization represents estimated state values. Eligibility traces extend this arrangement by maintaining one trace for each set of weights: a policy trace for the policy weights and a state-value trace for the state-value weights.

producesscalesscalesupdatesupdatesObserved transitionTD errorδPolicy traceeθPolicy weightsθState-value traceewState-value weightsw
How does one TD error flow through the actor-critic method to update policy weights and state-value weights?

w ← w + βδew

θ ← θ + αδeθ

Traces Through Earlier Visits

Eligibility traces preserve decayed information about earlier state visits and actions during an episode. A trace therefore carries some memory of what happened before the current step. The information from earlier visits is not retained at full strength indefinitely: it decays as the episode progresses. In an actor-critic method, the state-value trace and policy trace preserve this history for their respective weight updates.

episode advancesepisode advancesFirst visitpolicy and state-valuetraces record informationLater stepearlier information isdecayedCurrent steptraces carry earlierinformation into updates
As an episode progresses, how do traces retain information about earlier visits while that information decays?

Following One Earlier Visit

Consider an episode in which a state and an action occur, and later another step produces a TD error.

Earlier visit: The policy trace and state-value trace record information associated with the earlier state visit and action.

Progress to a later step: As the episode continues, the earlier information remains available through the traces but is decayed.

Weight update: When the later TD error is used, the traces carry information from the earlier visit into the policy-weight and state-value-weight updates.

Eligibility traces let a later TD error affect updates using information from earlier state visits and actions.

Policy and Value Parameterizations

Part of the methodWhat it representsTraceWeight update
Policy parameterizationThe policyPolicy trace using the gradient of the log policyθ ← θ + αδeθ
State-value parameterizationEstimated state valuesState-value trace using the gradient of the estimated state-value functionw ← w + βδew

The policy and state-value parts should not be treated as interchangeable. The state-value trace uses the gradient of the estimated state-value function. The policy trace uses the gradient of the log policy. These different gradients reflect the different parameterizations. Both updates use the same TD error, but each update applies that error through the trace belonging to its own parameterization.

formsupdatesformsupdatesscales updatescales updatePolicyparameterizationgradient of log policyPolicy traceeθPolicy weightsθState-valueparameterizationgradient of estimated statevalueState-value traceewState-value weightswTD errorshared learning signal
What does each parameterization represent, and how are they connected but updated differently?

Model-Free Incremental Learning

A method that requires no model can learn without being supplied a complete and accurate description of the environment. Temporal-difference learning is described as not requiring a model of the environment.

Fully incremental computation means that learning can proceed step by step. Temporal-difference learning can make progress from ongoing experience instead of waiting for a complete outcome before making progress.

processdrivesadvanceObserved transitionTD errorWeight updatespolicy and state-valueweightsNext stepincremental progress
How does TD learning update estimates from observed transitions without constructing or querying a model?

Choosing Among Three Method Classes

Dynamic programmingcomplete and accuratemodel; mathematically welldevelopedMonte Carlo methodsno model; not well suitedto incremental computationTemporal-differencelearningno model; fullyincremental; harder toanalyze
What information does each method require, when can it compute, and what trade-off does it make?
Method classModel requirementIncremental computationMain strengthMain weakness or trade-off
Dynamic programmingRequires a complete and accurate modelNot identified here as the defining featureMathematically well developedDepends on having the required model
Monte Carlo methodsDoes not require a modelNot well suited to step-by-step incremental computationConceptually simpleLess suited to incremental computation
Temporal-difference learningDoes not require a modelFully incrementalCan learn step by step without a modelMore complex to analyze

The comparison should not be reduced to a claim that one method is always fastest or best. The source notes that the methods differ in efficiency and speed of convergence but does not give a universal ranking. The clearest distinction for temporal-difference learning is its combination of no model requirement and fully incremental computation, balanced against greater analytical complexity.

Auditing an Update

MEDIUM

A learner reports that a later TD error changed both the policy weights and the state-value weights. Explain the four-part audit you would perform before deciding whether the update is correct.

Hints
  • Begin with the observed transition.
  • Then inspect the TD error and the two trace updates.
  • Finish by checking each weight update against its corresponding trace and parameter set.

A Four-Part Debugging Order

Determine the reliable order for inspecting an actor-critic update with eligibility traces.

1. Transition: Inspect the observed transition that supplied the current step's information.

2. TD error: Inspect the temporal-difference error computed from that step.

3. Trace updates: Inspect the state-value and policy eligibility traces, including the decayed information they carry from earlier visits and actions.

4. Weight updates: Inspect the state-value update using w ← w + βδew and the policy update using θ ← θ + αδeθ.

The reliable debugging order is transition, TD error, trace updates, and weight updates.

Common Update Mistakes

  • Treating the policy trace and state-value trace as one shared trace.

    The actor-critic method maintains one trace for the policy weights and another for the state-value weights.

    Fix: Track the policy trace and state-value trace separately.

  • Using the same gradient description for both traces.

    The state-value trace uses that gradient, while the policy trace uses the gradient of the log policy.

    Fix: Match each trace to its own parameterization and gradient.

  • Applying the state-value update to the policy weights, or the policy update to the state-value weights.

    The two equations update different weight sets.

    Fix: Use w ← w + βδew for state-value weights and θ ← θ + αδeθ for policy weights.

  • Assuming that no model means no observed experience is needed.

    The defining claim is that TD learning does not require a model; its process still proceeds from observed steps.

    Fix: Distinguish learning from experience from learning through a supplied model.

  • Assuming that a no-model method is automatically fully incremental.

    Monte Carlo methods do not require a model but are not well suited to step-by-step incremental computation.

    Fix: Check model requirement and incremental computation as separate properties.

When debugging, inspect the transition, TD error, trace updates, and weight updates in that order. This order follows the flow of information through the method and helps separate an incorrect learning signal from an incorrect trace or weight update.

Key Takeaways

  1. Eligibility traces preserve decayed information about earlier state visits and actions during an episode.
  2. An actor-critic method maintains separate policy and state-value traces.
  3. The state-value trace uses the gradient of the estimated state-value function, while the policy trace uses the gradient of the log policy.
  4. The TD error scales updates to both parameter sets: w ← w + βδew and θ ← θ + αδeθ.
  5. Temporal-difference learning requires no model and is fully incremental, while dynamic programming requires a complete and accurate model and Monte Carlo methods are not well suited to step-by-step incremental computation.

Key Takeaways

  • Eligibility traces carry decayed information from earlier state visits and actions into later updates.
  • The policy and state-value parameterizations have different roles, traces, gradients, and weight updates.
  • The TD error is the shared signal that scales both policy-weight and state-value-weight changes.
  • Temporal-difference learning is model-free and fully incremental.
  • Dynamic programming, Monte Carlo methods, and TD learning trade off model requirements, incremental computation, conceptual simplicity, and analytical complexity.