Concepts / State Values and Action Values

State Values and Action Values

A transition graph helps organize how an MDP moves from a state and action to possible next states and rewards.

  • Programming

A Decision Is Only the Start

When an agent chooses an action, the model must describe more than the choice itself. It must also describe what the environment may produce afterward: a new situation and a reward. The Recycling Robot makes this visible. It may search for a can, wait, or return to its home base to recharge. Each action can have different consequences, so the environment's response belongs in the model.

A transition graph organizes the possible one-step outcomes that follow a current state and a selected action.

Following One Transition

choosepossible outcomepossible outcomeCurrent stateSelected actionsearch, wait, or returnNext state Areward ANext state Breward B
From a current state and a selected action, what possible next states and rewards can occur?

Read the graph from left to right. Begin with the current state. Follow the selected action. Then inspect every branch leaving that action. Each branch represents a possible pair: a next state together with a reward. The graph therefore records the environment's possible one-step behavior rather than merely listing which actions are available.

What the Dynamics Record

For a chosen current state and action, the MDP records the probability of each possible pair consisting of a next state and a reward. The source describes these one-step dynamics as the complete description of a finite MDP's environment behavior. This is the most complete view in the supplied material because it preserves both which next state can occur and which reward accompanies that outcome.

can computecan computeOne-step dynamicsprobability of eachnext-state-and-reward pairState-transitionprobabilitiesprobability of movingbetween statesExpected rewardsexpected reward for astate-action pair
How do the full possible outcomes of a state-action pair relate to transition probabilities and expected rewards?
DescriptionWhat it preservesWhat it summarizes
Complete one-step dynamicsThe probability of every next-state-and-reward pairThe full one-step environment behavior
State-transition probabilitiesProbabilities of moving from one state to anotherThe state movement, without the complete pair detail
Expected rewardsThe reward expected for an action in a stateThe reward aspect for a state-action pair

State-transition probabilities and expected rewards are useful views, but each leaves out some detail compared with recording the probability of every next-state-and-reward pair. This distinction matters when two outcomes have different combinations of next states and rewards.

Tracing the Recycling Robot

choosechoosepossible outcomepossible outcomepossible outcomepossible outcomeDecision pointcurrent stateSearchNext state Areward AWaitNext state Breward BNext state Creward CNext state Dreward D
When the robot chooses an action such as searching or waiting, how can control move to different next states with different rewards?

Following a search decision

Trace what must be recorded after the Recycling Robot chooses to search.

Start with the current state: Identify the situation in which the robot reaches a decision point.

Record the selected action: The robot chooses search. The action is part of the one-step description.

List possible outcomes: Record every possible next state and the reward associated with that next state. The source emphasizes that one action may have multiple outcomes with different rewards.

Attach probabilities: For each next-state-and-reward pair, record its probability. This produces the complete one-step description for that state and action.

The result is not merely the label search. It is the set of possible next-state-and-reward pairs, together with the probability of each pair.

The same tracing method applies to waiting and returning to the home base to recharge. Start at the current state, follow the chosen action, and then record all possible next states and rewards. The action name alone does not tell us the environment's full response.

Values and the Supplied Model

The article title names state values and action values, but the supplied source material focuses on the transition graph and the environment's one-step dynamics. It explains expected rewards for state-action pairs, yet it does not provide formal numerical definitions or calculation rules for state values or action values. Therefore, the grounded distinction available here is between the current state, the selected action, and the possible next-state-and-reward outcomes that follow.

Model elementQuestion it organizesDetail supplied by the source
Current stateWhat situation is the robot in at the decision point?The starting point of a one-step transition
Selected actionWhat does the robot choose to do?Search, wait, or return to the home base to recharge
Next state and rewardWhat can the environment produce after the action?A possible pair recorded by the one-step dynamics
Expected rewardWhat reward is expected for this state-action pair?A useful quantity derived from the complete dynamics

Common Tracing Mistakes

  • Treating an action as if it had only one guaranteed outcome.

    The Recycling Robot example shows why one action may have multiple possible outcomes with different rewards.

    Fix: After naming the action, list every possible next-state-and-reward pair and its probability.

  • Recording only the next state.

    The complete one-step description records pairs consisting of a next state and a reward.

    Fix: Keep the next state and its associated reward together.

  • Calling state-transition probabilities the complete environment dynamics.

    State-transition probabilities are derived from the complete dynamics and leave out some detail.

    Fix: Use the complete next-state-and-reward probabilities when the full one-step behavior is required.

  • Confusing expected rewards with the full set of outcomes.

    Expected rewards are useful summaries, but they do not preserve every possible next-state-and-reward pair.

    Fix: Treat expected rewards as a derived view rather than a replacement for the complete dynamics.

Practice the Graph

MEDIUM

The Recycling Robot is at a decision point and chooses to wait. Describe what a complete one-step record must contain, without inventing particular state names, reward amounts, or probabilities.

Hints
  • Begin with the current state and the selected action.
  • List every possible next-state-and-reward pair.
  • Attach the probability of each pair.
  • Explain why an expected reward alone would not be the complete record.
  1. A strong answer identifies waiting as the selected action, then records all possible next states and their rewards together with the probability of each combined outcome. It also distinguishes that complete record from an expected reward, which is only a derived summary.

Key Takeaways

  • A transition graph begins with a current state and a selected action, then branches to possible next states and rewards.
  • The complete one-step dynamics record the probability of every possible next-state-and-reward pair.
  • State-transition probabilities and expected rewards are useful quantities derived from the complete dynamics, but each omits some detail.
  • The Recycling Robot may search, wait, or return to its home base, and one action may lead to multiple outcomes with different rewards.
  • The supplied material explains transition dynamics and expected rewards, but it does not provide formal numerical definitions for state values or action values.