Concepts / State Transitions and Environment Models

State Transitions and Environment Models

A cognitive map captures learned relationships among successive stimuli.

  • Programming

From Successions to Maps

Imagine an agent exploring an environment. It encounters one stimulus, then another, and later another. If these successions are learned, the first stimulus can become a cue for what is likely to appear next. A cognitive map therefore captures learned relationships among successive stimuli, rather than merely storing unrelated observations.

followed byfollowed bycontributes tocontributes tocontributes toStimulus S1observed firstCognitive maplearned relationshipsStimulus S2observed nextStimulus S3observed later
How do successive stimuli become connected into a map of relationships between states?

Tracing an Expectation

A Learned Sequence

Suppose an agent repeatedly encounters a generated sequence of three stimuli: a marked path, then a doorway, then a room.

Observe the first succession: The agent experiences the marked path followed by the doorway. This supplies a stimulus-stimulus relationship: the first stimulus is associated with what follows it.

Observe the next succession: The agent experiences the doorway followed by the room. This adds another relationship to the learned description.

Generate an expectation: After learning the succession, observing the marked path can activate an expectation that the doorway will occur next. The expectation comes from the learned sequence, not from a reward-only record.

Extend the map: The doorway can also point toward the room, so the learned description represents connected relationships among successive stimuli.

The successive observations support a cognitive-like map in which one observed stimulus points toward a subsequent stimulus.

expectationfollowed byStimulus Acurrent observationStimulus Bexpected nextStimulus Csubsequent observation
How does one stimulus activate an expectation of which stimulus will occur next?

Expectancy theory explains how a current stimulus can generate an expectation about what comes next. In the simplest stimulus-stimulus association, written here as S-S', S is the current state or stimulus and S' is the subsequent state or stimulus. After learning from such labeled successions, observing S leads the model to generate the expectation that S' will be observed next.

Inputs, Outputs, and System Models

System identification provides a control-engineering parallel. The learner observes inputs and labeled outputs and uses those transitions to learn a model of what the system will do next. In the simplest discrete-time case, an S-S' pair supplies a current state and its subsequent state as the label. When actions matter, the example becomes SA-S': the expected next state depends on both the current state and the executed action.

labeled transitionlabels transitionpredictsObserved inputstate or state-action pairTransition modellearned relationshipsExpected next stategenerated from the modelObserved outputsubsequent state
How does observing inputs and outputs allow a system to infer the environment's state-transition model?

Keep two questions separate when analyzing a learned environment model: What tends to follow the current state, and what is expected after a particular action in that state? The first corresponds to S-S'; the second corresponds to SA-S'. This distinction prevents an action-dependent transition from being mistaken for a transition that occurs without a choice by the agent.

States, Actions, and Rewards

The notation identifies what each training example is trying to learn. S refers to a state or stimulus, A refers to an action, S' refers to the expected subsequent state, and R refers to a reward signal. S-S' describes succession without including an agent choice. SA-S' adds the action so that the expected next state depends on both state and action. SR associates a reward with a state, while SAR associates a reward with a state-action pair.

Training formWhat is representedQuestion it supports
S-S'A state or stimulus and its subsequent stateWhat tends to follow this state?
SA-S'A state, an action, and the subsequent stateWhat is expected after taking this action in this state?
SRA state and an associated reward signalWhat reward signal is associated with this state?
SARA state, an action, and an associated reward signalWhat reward signal is associated with this state-action pair?

The forms distinguish transition information from reward information.

paired withassociated with transition toSR associationSAR association with state-action pairState Scurrent situationAction Achoice by agentState S'subsequent situationReward Rreward signal
What is the difference between the state, the action taken, and the reward received in one training example?

Learning Without Reward

A cognitive-like map does not require learning reward information at the same time as learning environmental succession. The source describes S-S', SA-S', SR, and SAR as forms of supervised learning that can help an agent acquire cognitive-like maps. In particular, an agent can explore, receive no non-zero reward signals, and still learn from labeled successive states. The succession information remains available even when reward information contributes nothing non-zero.

trainsdoes not preventcan produceLabeled S-S'examplessuccessive statesSupervised learninglearned associationsCognitive-like maprelationships among statesZero reward signalsno non-zero reward
How can labeled successive states produce a map of relationships even when every reward signal is zero?

Common Reasoning Errors

  • Assuming a cognitive map is only a record of rewards.

    The source describes cognitive-like maps as learnable from successive stimuli and states, separately from reward information.

    Fix: Ask first whether the learner is modeling state succession. Reward associations are additional information, not a required explanation for every learned relationship.

  • Leaving the action out of an action-dependent transition.

    SA-S' includes both the current state and the executed action.

    Fix: Represent the input as the state-action pair when the question is what happens after taking a particular action in a state.

  • Confusing S-S' with SR.

    S-S' represents a current state and a subsequent state, whereas SR associates a reward signal with a state.

    Fix: Check whether the label is another state or a reward signal before interpreting the example.

  • Treating system identification as unrelated to expectancy.

    The parallel is that observed inputs and labeled outcomes support learning what the system will do next.

    Fix: Connect the learned transition model to the expectation of a subsequent state.

Practice and Summary

EASY

Classify each description as S-S', SA-S', SR, or SAR. First: a current state is paired with the state observed next. Second: a current state and an executed action are paired with the state observed next. Third: a state is paired with a reward signal. Fourth: a state and action are paired with a reward signal.

Hints
  • Look for a subsequent state label, written S'.
  • Look for an action A before deciding whether the example is action-dependent.
  • Look for a reward signal R rather than a subsequent state.
  1. Successive stimuli can become a cognitive map because learned relationships connect what is observed now with what follows. Expectancy theory describes how a current stimulus can activate an expectation of a subsequent stimulus. System identification offers a control-engineering parallel in which labeled inputs and outputs support a learned transition model. S-S' describes state succession, SA-S' adds an action, SR associates reward with a state, and SAR associates reward with a state-action pair. Supervised learning from labeled state transitions can produce cognitive-like maps even when reward signals are all zero.

Key Takeaways

  • A cognitive map captures learned relationships among successive stimuli or states.
  • Expectancy theory explains how observing a current stimulus can generate an expectation about what comes next.
  • System identification learns a transition model from labeled inputs and outcomes.
  • SA-S' includes actions in state-transition learning, while SR and SAR include reward information.
  • Cognitive-like maps can be learned from supervised state-transition examples without non-zero reward signals.