Concepts / Temporal-Difference Learning and US Predictions

Temporal-Difference Learning and US Predictions

A TD model's stimulus representation is a substantive modeling choice, not a minor implementation detail.

  • Programming

The Representation Is Part of the Model

A temporal-difference model does not receive an external stimulus in a single, unavoidable format. It must translate what happens during a trial into state features that can support predictions. That translation is a substantive modeling choice: predictions depend critically on the stimulus representation, not only on learning parameters or on the mechanism used to generate a response.

The central question is how much the representation distinguishes one moment of a stimulus from another.

Suppose one component conditioned stimulus is presented for several moments. One representation can preserve the order of those moments, another can treat every moment of the presentation as the same feature, and a third can connect nearby moments through overlapping activity. These alternatives create a temporal-generalization gradient from no forced sharing to complete sharing.

Three Ways to Encode Time

separate featureseparate featureshared featureshared featurepartial overlappartial overlapt1CSC feature 1t1component featuret1microstimulit2CSC feature 2t2component featuret2overlapping microstimulit3CSC feature 3t3component featuret3overlapping microstimuli
How does the same external stimulus get encoded across time under CSC, MS, and presence representations?

The CSC representation assigns separate features to distinct time points during the presentation of a component stimulus. Its sequence begins at stimulus onset and continues with precisely timed, short-duration internal signals while the external stimulus is present. Nearby moments therefore do not automatically share a feature. CSC is useful when the model should have access to a precise temporal sequence, but CSC is not an essential part of every TD model.

The presence representation creates one feature for each component stimulus. That feature has value 1 whenever the component is present and 0 when it is absent. If the component lasts across several time steps, every one of those steps uses the same component feature. This representation therefore compresses the most information about time.

The MS representation lies between CSC and presence. A stimulus starts a cascade of microstimuli, and these are extended rather than brief, nonoverlapping pulses. Several microstimuli can be active at once, so nearby moments share some activity without becoming identical. Later microstimuli can become progressively wider and reach lower maximum activity, depending on the particular MS representation. MS is therefore a family of possible representations rather than one fixed pattern.

RepresentationEncoding across a component presentationSharing between nearby momentsTemporal generalization
CSCSeparate features for distinct time pointsNo forced feature sharingLowest
MSOverlapping microstimuliPartial sharing through overlapIntermediate
PresenceOne component feature while presentAll present moments share that featureHighest

The representations differ in how strongly they generalize learning across nearby time points.

Temporal Generalization and US Timing

Temporal generalization is the extent to which learning at one moment carries over to nearby moments. It determines how specifically a model can learn when an unconditioned stimulus, or US, is predicted. A representation that keeps moments distinct permits finer timing. A representation that shares features across moments encourages a prediction that is broad across the shared interval.

distinct prediction supportdistinct prediction supportdistinct prediction supportshared prediction supportt1separate featurecomponent presentshared featureUS predictionfine timingt2separate featureUS predictionbroad timingt3separate feature
When a stimulus feature generalizes across multiple time points, how does that change the timing precision of the predicted US?

Choosing between presence and MS

Two models represent one component stimulus that remains present for several time steps. The first uses presence; the second uses MS. Which model gives nearby moments completely shared stimulus support, and which gives partial temporal generalization?

Identify the first representation: Presence creates one feature for the component and uses it at every time step while the component is present. Nearby moments therefore receive completely shared stimulus support.

Identify the second representation: MS creates extended microstimuli. Several can be active at once, so nearby moments are connected by overlapping activity rather than by one identical component feature.

Relate representation to prediction timing: The presence model is pushed toward a more time-general US prediction. The MS model can preserve some timing distinction because its overlap is partial rather than complete.

Presence produces complete temporal generalization across the component presentation, whereas MS produces partial temporal generalization.

TD Learning in a Finite Decision Problem

Temporal-difference learning is a class of methods for solving finite Markov decision problems. The learner encounters states, takes actions, observes transitions and rewards, and maintains value estimates. Its defining computational feature is that learning can proceed incrementally, step by step, rather than waiting for a complete outcome before making progress.

chooseproducesrevealsupdatesupdatessupports next stepStatecurrent situationActionchoiceTransitionnext stateRewardobserved outcomeValue estimateupdated prediction
How do states, actions, transitions, rewards, and value estimates connect as TD learning solves a finite Markov decision problem?

The phrase finite Markov decision problem identifies the problem class being addressed; it does not specify one unique stimulus representation. A TD implementation still needs a way to represent the states or stimulus moments that support its predictions. Thus, a model can use TD learning while making different assumptions about temporal representation.

No Model and Incremental Updates

When a method requires no model, it does not depend on being supplied with a complete and accurate description of the environment's future transitions and rewards. Temporal-difference learning can use observed experience directly while learning. This is the sense in which TD learning is model-free in the comparison provided here.

provides evidenceprovides estimatechangesObserved experiencestate, transition, rewardCurrent valueexisting estimateTD updatestep-by-step adjustmentUpdated valuenew estimate
How can TD learning update predictions from observed transitions and rewards without storing or using a model of future transitions and rewards?

Fully incremental computation means that learning can make progress after each new step of experience. The method does not have to wait for a complete episode or a complete outcome before adjusting its value estimates. Each observed transition and reward can contribute to the next update, so information moves through the learning process one step at a time.

updatecontinueupdatecontinueupdateExperience 1transition and rewardEstimate 1updated valueExperience 2next transition and rewardEstimate 2updated valueExperience 3next transition and rewardEstimate 3updated value
What changes in the value estimate immediately after each new experience, and how does information move through the step-by-step TD update?

Three Method Classes Compared

requireswaits forusesDynamic programmingmodel required; incrementalstatus not definingComplete accuratemodelrequired by DPMonte Carlono model; complete outcomeComplete outcomeused by MCTemporal differenceno model; step by stepObserved stepused by TD
What information does each method require, when does each method update its estimates, and what are the main strengths and weaknesses?
MethodModel requirementWhen progress can be madeMain strengthMain weakness
Dynamic programmingRequires a complete and accurate modelBased on the supplied modelMathematically well developedDepends on having the required model
Monte CarloDoes not require a modelNot well suited to step-by-step incremental computationConceptually simpleCannot make progress in the same fully incremental way as TD methods
Temporal-difference learningDoes not require a modelStep by step from observed experienceCombines no model requirement with fully incremental computationMore complex to analyze

The comparison concerns information requirements and update style; the source does not give a universal speed or efficiency ranking.

The important comparison is not that one method is always best. Dynamic programming is mathematically well developed but requires a complete and accurate model. Monte Carlo methods do not require a model and are conceptually simple, but they are not well suited to step-by-step incremental computation. TD methods require no model and are fully incremental, although they are more complex to analyze.

Mistakes in Interpretation

  • Treating CSC as the definition of TD learning

    CSC is one stimulus representation. TD learning is a class of methods for solving finite Markov decision problems, and the representation is a separate modeling choice.

    Fix: Describe CSC, MS, and presence as alternative assumptions about how the external stimulus becomes state features.

  • Assuming that all representations preserve the same timing information

    Presence uses one component feature at every moment while the component is present.

    Fix: Check whether nearby moments receive separate features, overlapping microstimuli, or one shared component feature.

  • Calling MS identical to either CSC or presence

    MS uses extended microstimuli, and several can be active at once. Its overlap creates an intermediate pattern of temporal generalization.

    Fix: Treat MS as a family of representations with partial overlap between nearby moments.

  • Interpreting model-free as learning without observed experience

    TD learning uses observed experience while avoiding dependence on a complete and accurate environmental model.

    Fix: Explain model-free learning as updating from observed states, transitions, and rewards rather than from a supplied complete model.

  • Claiming that TD learning must wait for a complete outcome

    Fully incremental computation means that TD learning can make progress step by step.

    Fix: Emphasize that new experience can contribute to value-estimate updates as learning proceeds.

Apply the Distinctions

MEDIUM

A model represents a component stimulus that lasts for several time steps. Predict how nearby time points are encoded under CSC, MS, and presence. Then explain which representation would place the strongest pressure toward a broad US prediction and which would support the finest timing distinction.

Hints
  • Ask whether the representation assigns separate features, overlapping microstimuli, or one shared component feature.
  • More sharing across time means more temporal generalization.
  • Use the representation choice and the TD method as separate parts of your explanation.
MEDIUM

Compare dynamic programming, Monte Carlo methods, and TD learning using two questions: Does the method require a complete and accurate model? Can it make progress fully incrementally, step by step? State one strength and one weakness for each method.

Hints
  • Dynamic programming requires a complete and accurate model.
  • Monte Carlo methods do not require a model but are not well suited to step-by-step incremental computation.
  • TD methods require no model and are fully incremental, but they are more complex to analyze.

Key Takeaways

  1. CSC separates nearby moments, presence treats all moments during a component presentation as one shared feature, and MS connects nearby moments through overlapping microstimuli.
  2. These representations create different levels of temporal generalization, which affects the timing precision of US predictions.
  3. CSC, MS, and presence are modeling assumptions about stimulus representation, not different definitions of TD learning.
  4. Temporal-difference learning solves finite Markov decision problems without requiring a complete and accurate model and can update estimates fully incrementally.
  5. Dynamic programming, Monte Carlo, and TD methods trade off model requirements, update timing, conceptual or analytical simplicity, and incremental computation.

Key Takeaways

  • Stimulus representation is a substantive part of a TD model.
  • CSC gives distinct temporal features, presence gives one shared feature across a component presentation, and MS provides intermediate overlap.
  • Temporal generalization controls how precisely US predictions can be tied to particular moments.
  • TD learning is model-free in the sense that it does not require a complete and accurate environmental model, and it is fully incremental because it can learn step by step.
  • Dynamic programming, Monte Carlo, and TD learning differ in their information requirements and update styles; none is universally best on the evidence provided.