Concepts / Model Availability

Model Availability

An action value evaluates a state together with a particular action.

  • Programming

From Situation to Decision

Knowing that a situation is valuable does not automatically tell an agent what to do next. The missing information may be the effect of each available action. The central distinction in this topic is between evaluating a state by itself and evaluating a state together with a particular action.

A state value answers, in effect, how valuable a situation is. An action value evaluates a state-action pair: the current state together with one particular action.

combinecombineevaluateStatecurrent situationState-action pairevaluated togetherAction valuevalue of the pairActionchosen possibility
How do the current state and a particular action combine to determine an action value?

One-Step Look-Ahead

When a model is available, state values can be used in a one-step look-ahead decision. The agent considers an available action, uses the model to determine what that action will do, and then uses the value of the possible next state when comparing actions. In this setting, state values can be enough to determine which action the policy should select because the model connects each candidate action to its possible next state.

considerpredict effectsaction a1action a2compare valuecompare valueCurrent statesAvailable actionsa1, a2Modelaction effectsNext state 1value availablePolicy choiceselected actionNext state 2value available
How can the model use the value of each possible next state to choose the best action from the current state?

Comparing two actions through their next states

An agent is in state S and can choose action A or action B. A model is available. The model indicates that A leads to state X and B leads to state Y. The known state value of X is higher than the known state value of Y.

Start with the current state: The agent begins in S and lists the actions available there.

Use the model: The model connects action A with X and action B with Y.

Use state values: The agent compares the value of X with the value of Y.

Select an action: Because X has the higher state value in this illustration, the comparison favors action A.

With a model available, state values can support a policy decision through one-step look-ahead.

When the Model Is Missing

When no model is available, knowing the value of the current state does not identify the quality of the actions available from that state. The agent may know that the current situation is valuable, but it does not know what each action will do. Without that action-to-outcome connection, the state value alone cannot determine which action should be selected.

value comparisonone-step look-aheadinsufficient aloneunknown connectionCurrent state valueknownPolicy choicesupported by look-aheadCurrent state valueknownAction qualitynot identified by statevalue aloneModelconnects actions tooutcomesAction effectsnot identified
If only the value of the current state is known, how can the agent choose among actions when it cannot predict each action's next state and reward?

Why one state value cannot rank actions

An agent knows that state S has a high value. From S, it can choose action A or action B, but no model is available to indicate what either action will do.

Record what is known: The value of S describes the situation as a whole.

Identify what is missing: There is no information connecting A or B to their possible outcomes.

Check whether the actions can be ranked: The single value for S does not say whether A or B is the better action.

The state value alone is insufficient for choosing between A and B when the model is unavailable.

  • Treating the value of the current state as the value of every available action.

    A state value evaluates the situation, while an action value evaluates the situation together with a particular action.

    Fix: Ask whether the information distinguishes the available state-action pairs.

  • Assuming that state values can always determine a policy.

    The one-step look-ahead use of state values depends on having a model that connects actions with possible next states.

    Fix: Separate the model-available case from the model-unavailable case.

  • Ignoring action identity when evaluating a state-action pair.

    An action value requires both the state and a particular action.

    Fix: Represent the evaluation as a pair: current state plus selected action.

Action Values for Policy Suggestions

Estimating each available action's value makes the information useful for suggesting a policy. Instead of having only one value attached to the current state, the agent can compare the values associated with the different state-action pairs for that state. This comparison provides a direct basis for suggesting which action to take.

evaluatecomparecompareState Sone state valueState valuedoes not separate actionsAction Avalue for S with APolicy suggestioncompare action valuesAction Bvalue for S with B
How does comparing action values for the same state identify which action the agent should choose?

Comparing action values in one state

For the same state S, suppose an estimate is available for the pair consisting of S and action A, and another estimate is available for the pair consisting of S and action B. The estimate for the pair with A is higher.

Hold the state fixed: Both evaluations concern S, so the comparison is about different actions from the same situation.

Compare the action values: The estimate for the pair with A is compared with the estimate for the pair with B.

Suggest a policy choice: The higher estimated action value provides the stronger suggestion for the action to select.

Estimating action values turns action selection into a comparison among state-action pairs.

EvaluationWhat it includesUse in this topic
State valueA state or situationCan support action choice through one-step look-ahead when a model is available
Action valueA state together with a particular actionAllows action values for the same state to be compared directly

Monte Carlo and q*

Monte Carlo methods connect with this topic through the goal of estimating q*. The relevant target is not only the value of a state, but the value associated with each state-action pair. Sampled episodes provide returns for those pairs, and repeated returns can improve the estimates used to compare actions. In this way, Monte Carlo estimation supports policy suggestions even when a model is not available.

observeuse returnrepeatupdatecontinue samplingState-action pairS with action AEpisode samplefirst observed returnq* estimatecurrent estimateMore episodesadditional returnsRefined estimateupdated comparison
How do sampled episodes produce returns for state-action pairs, and how do repeated returns improve estimates of q*?

What do you think happens?

An agent has no model, but it has estimates for the action values of two actions from the same state. Which information is more directly useful for suggesting the next action: one value for the state, or the two action-value estimates?

  • One value for the state
  • The two action-value estimates
  • Neither can support a suggestion
Reveal answer

Answer: The two action-value estimates

Action values evaluate the state together with particular actions, so comparing them directly distinguishes the available choices. This is the role connected with estimating q*.

Practice and Transfer

MEDIUM

A learner says: “I know the current state has a high value, so I know which action to take.” Explain why this conclusion depends on model availability. Then describe what information would be more useful when no model is available.

Hints
  • Ask whether the state value identifies what each action will do.
  • Distinguish a state evaluation from a state-action evaluation.
  • Explain how comparing estimated action values could suggest a policy.

When analyzing an action-selection problem, use this order: identify whether a model is available; determine whether the information evaluates states or state-action pairs; then ask whether the available values can be compared across the actions from the current state.

  1. Model availability changes how state values can be used. With a model, state values support one-step look-ahead because the model connects actions to possible next states. Without a model, the current state's value does not identify action quality by itself. Action values solve the comparison problem by evaluating each state together with a particular action. Estimating those values supports policy suggestions and connects Monte Carlo methods with the goal of estimating q*.

Key Takeaways

  • An action value evaluates a state together with a particular action.
  • A model allows state values to support a one-step look-ahead policy decision.
  • Without a model, a state value alone does not identify which available action is better.
  • Estimating action values makes it possible to compare actions from the same state.
  • Monte Carlo methods connect with the goal of estimating q* for state-action pairs.