State Transitions and Environment Models
A cognitive map captures learned relationships among successive stimuli.
From Successions to Maps
Imagine an agent exploring an environment. It encounters one stimulus, then another, and later another. If these successions are learned, the first stimulus can become a cue for what is likely to appear next. A cognitive map therefore captures learned relationships among successive stimuli, rather than merely storing unrelated observations.
Tracing an Expectation
A Learned Sequence
Suppose an agent repeatedly encounters a generated sequence of three stimuli: a marked path, then a doorway, then a room.
Observe the first succession: The agent experiences the marked path followed by the doorway. This supplies a stimulus-stimulus relationship: the first stimulus is associated with what follows it.
Observe the next succession: The agent experiences the doorway followed by the room. This adds another relationship to the learned description.
Generate an expectation: After learning the succession, observing the marked path can activate an expectation that the doorway will occur next. The expectation comes from the learned sequence, not from a reward-only record.
Extend the map: The doorway can also point toward the room, so the learned description represents connected relationships among successive stimuli.
The successive observations support a cognitive-like map in which one observed stimulus points toward a subsequent stimulus.
Expectancy theory explains how a current stimulus can generate an expectation about what comes next. In the simplest stimulus-stimulus association, written here as S-S', S is the current state or stimulus and S' is the subsequent state or stimulus. After learning from such labeled successions, observing S leads the model to generate the expectation that S' will be observed next.
Inputs, Outputs, and System Models
System identification provides a control-engineering parallel. The learner observes inputs and labeled outputs and uses those transitions to learn a model of what the system will do next. In the simplest discrete-time case, an S-S' pair supplies a current state and its subsequent state as the label. When actions matter, the example becomes SA-S': the expected next state depends on both the current state and the executed action.
Keep two questions separate when analyzing a learned environment model: What tends to follow the current state, and what is expected after a particular action in that state? The first corresponds to S-S'; the second corresponds to SA-S'. This distinction prevents an action-dependent transition from being mistaken for a transition that occurs without a choice by the agent.
States, Actions, and Rewards
The notation identifies what each training example is trying to learn. S refers to a state or stimulus, A refers to an action, S' refers to the expected subsequent state, and R refers to a reward signal. S-S' describes succession without including an agent choice. SA-S' adds the action so that the expected next state depends on both state and action. SR associates a reward with a state, while SAR associates a reward with a state-action pair.
| Training form | What is represented | Question it supports |
|---|---|---|
| S-S' | A state or stimulus and its subsequent state | What tends to follow this state? |
| SA-S' | A state, an action, and the subsequent state | What is expected after taking this action in this state? |
| SR | A state and an associated reward signal | What reward signal is associated with this state? |
| SAR | A state, an action, and an associated reward signal | What reward signal is associated with this state-action pair? |
The forms distinguish transition information from reward information.
Learning Without Reward
A cognitive-like map does not require learning reward information at the same time as learning environmental succession. The source describes S-S', SA-S', SR, and SAR as forms of supervised learning that can help an agent acquire cognitive-like maps. In particular, an agent can explore, receive no non-zero reward signals, and still learn from labeled successive states. The succession information remains available even when reward information contributes nothing non-zero.
Common Reasoning Errors
Assuming a cognitive map is only a record of rewards.
The source describes cognitive-like maps as learnable from successive stimuli and states, separately from reward information.
Fix:
Ask first whether the learner is modeling state succession. Reward associations are additional information, not a required explanation for every learned relationship.Leaving the action out of an action-dependent transition.
SA-S' includes both the current state and the executed action.
Fix:
Represent the input as the state-action pair when the question is what happens after taking a particular action in a state.Confusing S-S' with SR.
S-S' represents a current state and a subsequent state, whereas SR associates a reward signal with a state.
Fix:
Check whether the label is another state or a reward signal before interpreting the example.Treating system identification as unrelated to expectancy.
The parallel is that observed inputs and labeled outcomes support learning what the system will do next.
Fix:
Connect the learned transition model to the expectation of a subsequent state.
Practice and Summary
Classify each description as S-S', SA-S', SR, or SAR. First: a current state is paired with the state observed next. Second: a current state and an executed action are paired with the state observed next. Third: a state is paired with a reward signal. Fourth: a state and action are paired with a reward signal.
Hints
- Look for a subsequent state label, written S'.
- Look for an action A before deciding whether the example is action-dependent.
- Look for a reward signal R rather than a subsequent state.
- Successive stimuli can become a cognitive map because learned relationships connect what is observed now with what follows. Expectancy theory describes how a current stimulus can activate an expectation of a subsequent stimulus. System identification offers a control-engineering parallel in which labeled inputs and outputs support a learned transition model. S-S' describes state succession, SA-S' adds an action, SR associates reward with a state, and SAR associates reward with a state-action pair. Supervised learning from labeled state transitions can produce cognitive-like maps even when reward signals are all zero.
Key Takeaways
- A cognitive map captures learned relationships among successive stimuli or states.
- Expectancy theory explains how observing a current stimulus can generate an expectation about what comes next.
- System identification learns a transition model from labeled inputs and outcomes.
- SA-S' includes actions in state-transition learning, while SR and SAR include reward information.
- Cognitive-like maps can be learned from supervised state-transition examples without non-zero reward signals.