Concepts / Action Selection from the Current State

Action Selection from the Current State

A state signal may combine immediate sensations with information derived from past sensations.

  • Programming

The Decision-Making Question

An agent does not always choose an action from the entire history of everything it has sensed. Instead, it can use a current state signal: the information currently available to the agent and relevant to deciding what to do next. The important question is not whether this signal contains every fact about the environment. The important question is whether it preserves the information needed to predict what may happen next, what rewards may be expected, and which action should be selected.

derived informationimmediate informationbasis forsupportsEarlier sensationsinformation derived fromthe pastCurrent stateinformation available nowAction selectionchoose an actionLatest sensationimmediate informationFuture predictionsstates and rewards
How can the current state combine the latest sensation with information remembered or derived from earlier sensations?

Tracing Information into State

A common first interpretation says that the state is whatever the agent senses right now. That interpretation is too narrow. A state may include immediate measurements, but it may also include information constructed from earlier sensations. For example, a system may recognize that an object remains present after it is no longer visible, or interpret the word yes differently depending on the earlier question. In each case, the latest sensation alone does not describe all the information currently available to the agent.

Latest sensation versus current state

An agent receives a latest sensation that does not identify an object clearly. Earlier sensations had already established that the object was present. What can the current state contain?

Read the latest sensation: The agent receives immediate information, but the latest sensation may not contain every fact that was previously available.

Use earlier information: The agent may construct or retain information from earlier sensations, such as the recognition that the object remains present after it is no longer visible.

Form the current state: The current state can combine the latest sensation with the relevant information derived from earlier sensations.

Select an action: The agent can base action selection on this richer state rather than on the latest raw sensation alone.

A state is the information available to the agent now, not necessarily a copy of the latest raw sensation.

preserve relevant informationrepresentpredictestimateselectSensation historypast and latestobservationsRetained informationdetails relevant to whathappens nextState representationcompact current informationFuture statespredictionExpected rewardspredictionSelected actionsdecision
What information from the history is preserved in the state, and how does that information affect predictions of future states, rewards, and actions?

The Markov Property

A state is Markov when it preserves all information relevant to future states, expected rewards, and action selection. Once the current Markov state and an action are known, the complete history is unnecessary for decision-making.

The Markov property is about relevant information, not total information. A Markov state does not need to preserve every detail of how the current situation arose. It needs to preserve the details that can affect what happens next or which action should be selected. Any part of the past that cannot affect those predictions or decisions can be discarded.

compress relevant informationselectsupports predictionaffectssupports predictionPast historycomplete sensation recordCurrent staterelevant informationretainedActionchosen from the stateFuture statepredicted consequenceExpected rewardpredicted consequence
If the current state is known, can the future state, reward, and action be considered without needing the full history of past sensations?

Testing Information Loss

To test a state representation, ask whether two different histories can produce the same current state while implying different future states, expected rewards, or useful actions. If they can, the state has merged histories that should remain distinguishable for decision-making. The omitted information may be something such as motion information, including velocity, but the decisive issue is always its effect on prediction or action selection.

Two histories, one state signal

Imagine that a state representation records an object's current position but omits its velocity. Two histories produce the same recorded position: in one history the object was moving toward a boundary, and in the other it was moving away. Does the recorded position alone preserve all decision-relevant information?

Compare the recorded states: The representation gives the agent the same current state in both cases because the recorded position is the same.

Compare the omitted information: The histories differ in motion information. One includes movement toward the boundary and the other includes movement away from it.

Ask whether the difference matters: If the difference in motion changes what happens next, the future cannot be predicted equally well from the recorded position alone.

Evaluate action selection: If different actions would be appropriate in the two situations, the position-only signal is a worse basis for action selection than a representation that retains the relevant motion information.

The representation is not fully Markov whenever the omitted motion information affects future states, expected rewards, or the action that should be selected.

preserves relevant distinctionssupportsmapped tomapped tocannot distinguishRelevant historyrelevant informationretainedDistinct currentstatessituations remaindistinguishableDecision-relevantpredictionfuture and rewardinformation preservedHistory Adifferent relevant pastSame current staterelevant distinctiondiscardedDifferentimplicationsfuture, reward, or actionmay differHistory Bdifferent relevant past
What changes when two different histories produce the same current state even though they imply different future states or rewards?

Useful Approximation

Real state signals may not satisfy the Markov property perfectly. That does not make them useless. A non-Markov signal can still work well when it is a good approximation to a Markov state. The closer the representation comes to preserving information needed for future rewards and decisions, the better the reinforcement learning system can be expected to perform.

supportssupports more effectivelyPractical statesome relevant informationomittedUseful actiondecision remains effectiveCloser Markovapproximationmore relevant informationretainedBetter-informedactionprediction and selectionimprove
Why can a state representation still support useful action selection even when it has discarded some information needed for perfect prediction?

Selecting from the State

combine and deriveuseinformSensationslatest and earlierinformationCurrent staterelevant informationavailable nowPredict consequencesfuture states and expectedrewardsChoose actiondecision based on currentstate
How does an agent use the current state signal to choose an action without directly processing the entire history of sensations?
MEDIUM

A state signal records only the latest raw sensation. Earlier sensations contained information that would affect which action is best, but that information is no longer available in the signal. Is the signal fully Markov? Explain your answer in terms of future states, expected rewards, and action selection.

Hints
  • Ask whether the omitted information could change what happens next.
  • Ask whether the omitted information could change the expected reward.
  • Ask whether two different histories could now require different actions.

What do you think happens?

Suppose the current state retains every detail from the past that can affect future states, expected rewards, or action selection. Do you still need the complete history to make the decision?

Reveal answer

Answer: No. Once the current state contains all relevant information, the complete history is unnecessary for decision-making.

This is the practical meaning of the Markov property: the state is a sufficient compact basis for prediction and action selection, even though it need not preserve every detail of how the situation arose.

Common Mistakes

  • Treating the state as only the latest raw sensation.

    A state may combine immediate measurements with information constructed from earlier sensations.

    Fix: Ask what information is currently available to the agent, including relevant information retained from the past.

  • Treating a Markov state as a complete description of everything in the environment.

    The Markov property does not require knowledge of hidden information that has never been revealed through sensations.

    Fix: Require preservation of relevant observed information, not omniscience.

  • Assuming that every detail of the past must be retained.

    A Markov state can be compact because irrelevant parts of the path do not matter once relevant information is present.

    Fix: Retain details that affect prediction or action selection and discard details that do not.

  • Declaring a state invalid merely because it is not perfectly Markov.

    A non-Markov signal can still be a useful approximation to a Markov state.

    Fix: Evaluate how well the representation supports prediction and action selection.

  • Naming a missing variable without checking whether it matters.

    The decisive issue is whether the omitted information makes prediction or action selection worse.

    Fix: Test whether the omitted information changes future states, expected rewards, or the appropriate action.

Key Takeaways

  1. A state signal can combine the latest sensation with information derived from earlier sensations.
  2. The Markov property means that the state preserves all information relevant to future states, expected rewards, and action selection.
  3. A Markov state does not need to contain hidden facts the agent has never observed or every irrelevant detail of the past.
  4. If relevant information is omitted, decisions based only on the current state may be less informed than decisions based on the complete history.
  5. A practical state can still be useful when it is a good approximation to a Markov state.

Key Takeaways

  • A state is the information available to the agent now, not necessarily the latest raw sensation.
  • A Markov state preserves information relevant to predicting future states, expected rewards, and action selection.
  • The agent need not know hidden information it has never observed, but it should not forget relevant information that was available.
  • State representations should be evaluated by how well they support prediction and decision-making.
  • Even an imperfect state signal can be useful when it closely approximates a Markov state.