Action Selection from the Current State
A state signal may combine immediate sensations with information derived from past sensations.
The Decision-Making Question
An agent does not always choose an action from the entire history of everything it has sensed. Instead, it can use a current state signal: the information currently available to the agent and relevant to deciding what to do next. The important question is not whether this signal contains every fact about the environment. The important question is whether it preserves the information needed to predict what may happen next, what rewards may be expected, and which action should be selected.
Tracing Information into State
A common first interpretation says that the state is whatever the agent senses right now. That interpretation is too narrow. A state may include immediate measurements, but it may also include information constructed from earlier sensations. For example, a system may recognize that an object remains present after it is no longer visible, or interpret the word yes differently depending on the earlier question. In each case, the latest sensation alone does not describe all the information currently available to the agent.
Latest sensation versus current state
An agent receives a latest sensation that does not identify an object clearly. Earlier sensations had already established that the object was present. What can the current state contain?
Read the latest sensation: The agent receives immediate information, but the latest sensation may not contain every fact that was previously available.
Use earlier information: The agent may construct or retain information from earlier sensations, such as the recognition that the object remains present after it is no longer visible.
Form the current state: The current state can combine the latest sensation with the relevant information derived from earlier sensations.
Select an action: The agent can base action selection on this richer state rather than on the latest raw sensation alone.
A state is the information available to the agent now, not necessarily a copy of the latest raw sensation.
The Markov Property
A state is Markov when it preserves all information relevant to future states, expected rewards, and action selection. Once the current Markov state and an action are known, the complete history is unnecessary for decision-making.
The Markov property is about relevant information, not total information. A Markov state does not need to preserve every detail of how the current situation arose. It needs to preserve the details that can affect what happens next or which action should be selected. Any part of the past that cannot affect those predictions or decisions can be discarded.
Testing Information Loss
To test a state representation, ask whether two different histories can produce the same current state while implying different future states, expected rewards, or useful actions. If they can, the state has merged histories that should remain distinguishable for decision-making. The omitted information may be something such as motion information, including velocity, but the decisive issue is always its effect on prediction or action selection.
Two histories, one state signal
Imagine that a state representation records an object's current position but omits its velocity. Two histories produce the same recorded position: in one history the object was moving toward a boundary, and in the other it was moving away. Does the recorded position alone preserve all decision-relevant information?
Compare the recorded states: The representation gives the agent the same current state in both cases because the recorded position is the same.
Compare the omitted information: The histories differ in motion information. One includes movement toward the boundary and the other includes movement away from it.
Ask whether the difference matters: If the difference in motion changes what happens next, the future cannot be predicted equally well from the recorded position alone.
Evaluate action selection: If different actions would be appropriate in the two situations, the position-only signal is a worse basis for action selection than a representation that retains the relevant motion information.
The representation is not fully Markov whenever the omitted motion information affects future states, expected rewards, or the action that should be selected.
Useful Approximation
Real state signals may not satisfy the Markov property perfectly. That does not make them useless. A non-Markov signal can still work well when it is a good approximation to a Markov state. The closer the representation comes to preserving information needed for future rewards and decisions, the better the reinforcement learning system can be expected to perform.
Selecting from the State
A state signal records only the latest raw sensation. Earlier sensations contained information that would affect which action is best, but that information is no longer available in the signal. Is the signal fully Markov? Explain your answer in terms of future states, expected rewards, and action selection.
Hints
- Ask whether the omitted information could change what happens next.
- Ask whether the omitted information could change the expected reward.
- Ask whether two different histories could now require different actions.
What do you think happens?
Suppose the current state retains every detail from the past that can affect future states, expected rewards, or action selection. Do you still need the complete history to make the decision?
Reveal answer
Answer: No. Once the current state contains all relevant information, the complete history is unnecessary for decision-making.
This is the practical meaning of the Markov property: the state is a sufficient compact basis for prediction and action selection, even though it need not preserve every detail of how the situation arose.
Common Mistakes
Treating the state as only the latest raw sensation.
A state may combine immediate measurements with information constructed from earlier sensations.
Fix:
Ask what information is currently available to the agent, including relevant information retained from the past.Treating a Markov state as a complete description of everything in the environment.
The Markov property does not require knowledge of hidden information that has never been revealed through sensations.
Fix:
Require preservation of relevant observed information, not omniscience.Assuming that every detail of the past must be retained.
A Markov state can be compact because irrelevant parts of the path do not matter once relevant information is present.
Fix:
Retain details that affect prediction or action selection and discard details that do not.Declaring a state invalid merely because it is not perfectly Markov.
A non-Markov signal can still be a useful approximation to a Markov state.
Fix:
Evaluate how well the representation supports prediction and action selection.Naming a missing variable without checking whether it matters.
The decisive issue is whether the omitted information makes prediction or action selection worse.
Fix:
Test whether the omitted information changes future states, expected rewards, or the appropriate action.
Key Takeaways
- A state signal can combine the latest sensation with information derived from earlier sensations.
- The Markov property means that the state preserves all information relevant to future states, expected rewards, and action selection.
- A Markov state does not need to contain hidden facts the agent has never observed or every irrelevant detail of the past.
- If relevant information is omitted, decisions based only on the current state may be less informed than decisions based on the complete history.
- A practical state can still be useful when it is a good approximation to a Markov state.
Key Takeaways
- A state is the information available to the agent now, not necessarily the latest raw sensation.
- A Markov state preserves information relevant to predicting future states, expected rewards, and action selection.
- The agent need not know hidden information it has never observed, but it should not forget relevant information that was available.
- State representations should be evaluated by how well they support prediction and decision-making.
- Even an imperfect state signal can be useful when it closely approximates a Markov state.