State Representation in Reinforcement Learning
A state signal may combine immediate sensations with information derived from past sensations.
The Information Available Now
In reinforcement learning, a state is not necessarily the latest raw sensation received by an agent. A state signal describes the information currently available to the agent. That information may include an immediate measurement, but it may also include information constructed from earlier sensations.
The central question is not whether the state contains everything in the environment. The question is whether it retains the information that matters for predicting what happens next, expected rewards, and which action should be selected.
What Markov Means
A state signal has the Markov property when it preserves all information relevant to future states, expected rewards, and action selection. Once the current state and action are known, the complete history of earlier sensations and actions is unnecessary for decision-making.
Markov does not mean that the state records every detail of how the current situation arose. A compact state may discard the path that led to it. That discarded path is harmless when the current state already contains every detail that can affect what happens next or which action should be chosen.
Building a State from History
The latest sensation can be ambiguous. Earlier sensations may provide context that changes the meaning of what is sensed now. A state representation can therefore combine the immediate sensation with information derived from previous sensations.
Context changes the current interpretation
An agent receives the word yes as its latest sensation. Can the latest sensation alone determine what the agent should understand?
Latest sensation: The immediate sensation is the word yes.
Earlier sensation: The earlier question gives the current word its context. The same yes can be interpreted differently depending on that earlier question.
State construction: A richer state can retain information derived from the earlier question together with the latest word.
Decision relevance: If the earlier question affects the action to select, forgetting it makes the current signal a weaker basis for action selection.
The state need not equal the latest raw sensation. It may include information derived from the earlier sensation when that information remains relevant.
The same idea applies when a system recognizes that an object remains present after it is no longer visible. The current state can contain information derived from earlier visual sensations rather than reporting only what is visible at this instant.
Testing Information Retention
To evaluate a representation, compare histories that it maps to the same current state. Ask whether those histories still differ in ways that matter for future states, expected rewards, or action selection. If relevant differences remain, the representation has discarded information that could improve prediction or decision-making.
| Question | What a sufficient representation does | Warning sign |
|---|---|---|
| Future states | Retains information relevant to predicting what happens next | Different histories with the same state lead to meaningfully different future predictions |
| Expected rewards | Retains information relevant to predicting rewards | The same represented state hides reward-relevant differences |
| Action selection | Provides as effective a basis for choosing actions as the complete history | A better action depends on history that the state omitted |
A Practical Sufficiency Check
- Identify two situations or histories that receive the same represented state.
- Ask whether their future-state predictions are the same for the decision being considered.
- Ask whether their expected rewards are the same or sufficiently similar for the task.
- Ask whether the same action choice remains appropriate in both situations.
- If a relevant difference was discarded, consider adding information derived from earlier sensations.
A missing variable matters only if its omission makes the state a worse basis for predicting what happens next or for selecting an action. Motion information such as velocity is a possible missing part of a state, but the decisive issue is its effect on prediction and decision quality, not the variable's name.
Useful Approximations
Real state signals may not satisfy the Markov property perfectly. That does not automatically make them useless. A non-Markov signal can still work well when it is a good approximation to a Markov state.
The practical standard is how well the representation supports prediction and action selection. The closer it comes to preserving information needed for future rewards and decisions, the better the reinforcement learning system can be expected to perform.
Mistakes in State Reasoning
Treating the state as exactly the latest raw sensation
Past sensations can contribute information that remains available to the agent and can affect action selection.
Fix:
Check whether the representation should include information derived from earlier sensations.Assuming a Markov state must describe everything in the environment
The Markov property concerns information available to the agent. It does not require knowledge of hidden information that was never revealed through sensations.
Fix:
Ask whether the state has forgotten relevant available information, rather than whether it knows every environmental fact.Preserving every detail of the full history
A Markov state can be compact. Details that cannot affect what happens next or what action should be selected need not be retained.
Fix:
Preserve relevant information, not history merely for its own sake.Declaring a representation useless because it is not perfectly Markov
A non-Markov signal can be a useful approximation to a Markov state.
Fix:
Evaluate how closely the representation preserves information needed for prediction and action selection.
Practice Check
A system receives a current sensation that is identical in two situations. In the first situation, earlier sensations contain information that would lead to one action. In the second, earlier sensations contain information that would lead to another action. The state representation maps both situations to the same state. Is the representation Markov for this decision?
Hints
- Compare the information retained by the shared state with the information in the two histories.
- Ask whether the discarded distinction affects action selection or future predictions.
- A representation can fail the Markov property when relevant available information has been forgotten.
Practice answer
Determine whether the shared state is Markov for the decision.
Compare the histories: The histories differ in earlier sensations even though their latest sensation is identical.
Check relevance: The earlier sensations affect which action should be selected, so the difference is relevant.
Check retention: The representation maps both histories to the same state and therefore does not retain that relevant distinction.
The representation is not Markov for this decision because the current state is not sufficient for action selection without the omitted history.
Key Takeaways
- A state signal is Markov when it preserves information relevant to future states, expected rewards, and action selection.
- A state may combine an immediate sensation with information derived from earlier sensations.
- A Markov state does not need hidden information the agent has never observed or every detail of the path that produced the current situation.
- To test a representation, compare situations mapped to the same state and check whether their future predictions, rewards, and action choices still differ.
- A non-Markov signal can still be useful when it is a good approximation to a Markov state.