Non-Markov State Signals
A state signal may combine immediate sensations with information derived from past sensations.
Why the Latest Sensation Is Not Enough
A state signal is the information available to an agent when it makes decisions. It may include what the agent senses immediately, but it may also include information constructed from earlier sensations. Therefore, the current state should not automatically be equated with the latest raw sensation.
An agent may look at a scene and build a richer representation. It may also recognize that an object remains present after it is no longer visible. In language, the word yes can be interpreted differently depending on the earlier question. In each case, earlier sensations influence what information is currently available.
Tracing Information Through Time
Consider an agent that first sees an object and later receives a sensation in which the object is not visible. If the current state contains only the latest raw sensation, the earlier sighting may no longer be available. If the state retains information derived from the earlier sighting, the current state can still represent that the object may remain present. The important question is not whether every past detail was preserved. The question is whether information relevant to what happens next or to the next action was preserved.
The Markov Test
A state is Markov when it preserves all information relevant to future states, expected rewards, and action selection. Once the current state and action are known, the complete history is unnecessary for decision-making.
To evaluate a state representation, ask whether two situations that look identical in the current state could nevertheless require different predictions or different actions because of information available earlier. If the omitted history contains information relevant to future states, rewards, or action selection, the current signal is not fully Markov. If the current state retains the relevant information, using that state can be as effective a basis for choosing actions as using the complete history.
Hidden Information and Forgotten Information
The Markov property does not require an agent to know everything about the environment. Hidden information can be useful without belonging to the agent's state when the agent has never received sensations that reveal it. The important failure is different: information that was available to the agent and is relevant to prediction or action should not be forgotten merely because it came from an earlier sensation.
A blackjack player is not expected to know the next card in the deck. Someone answering a phone is not expected to know the caller's identity in advance. A paramedic arriving at an accident is not expected to immediately know the internal injuries of an unconscious victim. These are hidden facts, not necessarily forgotten observations.
When judging a state signal, distinguish between information the agent never observed and relevant information the agent observed but failed to retain. Only the second case directly indicates that the representation may be losing useful history.
Predictive Relevance
Decisions and values are treated as functions of the current state. If the state omits relevant information, a decision based only on that state may be less informed than a decision based on the complete history. If the state retains the relevant information, the current state is an effective basis for predicting future states and rewards and for selecting actions.
Checking for Missing Motion Information
A state describes an object's current position but does not include information about its motion. How should you judge whether this is a useful state representation?
Identify the omitted information: Motion information, such as velocity, is a possible part of the situation that the state does not contain.
Test predictive value: Ask whether omitting that motion information makes the state a worse basis for predicting what happens next.
Test action selection: Ask whether an action selected from the current state would be less informed than an action selected using the complete history.
Judge the representation: The decisive issue is not the name of the missing variable. The issue is whether its omission makes future prediction or action selection worse.
A state representation is inadequate to the extent that omitted information is relevant to future states, rewards, or decisions.
Useful Approximation
Real state signals may not satisfy the Markov property perfectly. That does not make them automatically useless. A non-Markov signal can work well when it is a good approximation to a Markov state. The closer the representation comes to preserving the information needed for future rewards and decisions, the better the reinforcement learning system can be expected to perform.
Common Evaluation Mistakes
Treating the state as only the latest raw sensation.
A state may include information derived from earlier sensations, such as information that an object may remain present.
Fix:
Check whether the current state incorporates relevant information constructed from the observation history.Assuming a Markov state must contain everything in the environment.
The agent is not expected to know hidden information it has never observed.
Fix:
Focus on whether relevant information available to the agent was retained.Assuming every detail of the past must be preserved.
A Markov state can be compact and discard path details that are irrelevant once the relevant information is present.
Fix:
Ask whether the omitted detail affects future states, expected rewards, or action selection.Rejecting a practical signal because it is not perfectly Markov.
A non-Markov signal can still be a useful approximation to a Markov state.
Fix:
Evaluate how closely the representation preserves information needed for prediction and decisions.
Practice Check
A state signal records the latest sensation but omits information about an object's earlier motion. Explain how you would determine whether the signal is Markov. Your answer should discuss future-state prediction, expected rewards, action selection, and the difference between hidden information and forgotten information.
Hints
- Ask whether the omitted motion information changes what can happen next.
- Ask whether an action based only on the current signal would be less informed than one based on the complete history.
- Do not assume that every unobserved fact must be included.
What do you think happens?
If a state omits an earlier observation, does that automatically make the state non-Markov?
Reveal answer
Answer: No, only omitted information relevant to future states, expected rewards, or action selection makes it non-Markov.
A Markov state may discard details about how the current situation arose when those details cannot affect what happens next or which action should be selected.
Key Takeaways
- A state is Markov when it preserves all information relevant to future states, expected rewards, and action selection.
- The current state may combine the latest sensation with information derived from earlier sensations.
- A Markov state does not need to contain hidden information the agent has never observed or every irrelevant detail of the past.
- The critical test is whether forgotten information makes prediction or action selection worse.
- A non-Markov state signal can still be useful when it is a good approximation to a Markov state.
Key Takeaways
- The Markov property concerns preservation of information relevant to future states, rewards, and actions.
- A state can incorporate information derived from past sensations rather than representing only the latest sensation.
- Hidden facts the agent never observed are different from relevant observations that the state representation forgot.
- A practical state signal may be valuable even when it is only approximately Markov.