State-Action Value Functions
The Bellman equation is a recursive consistency relationship for value functions.
Value Beyond the Current State
A value function does not need to evaluate every state as an isolated case. Under a policy π, the value of a state s is connected to what may happen after the agent is in s. The agent may receive a reward, move to a successor state, and then receive the value associated with that successor. The Bellman equation expresses this recursive consistency relationship.
The central question is: how should the value of the current state reflect immediate rewards and the values of the states that may follow it?
Reading a Backup Relationship
A backup relationship moves value information from possible successor states back toward the state currently being evaluated. Begin with a state. The policy determines possible actions, the environment produces rewards and successor states, and the values of those successor states contribute to the value of the starting state. The word backup refers to this movement of information from later possibilities toward the current estimate.
In a backup diagram, open circles represent states and solid circles represent state-action pairs. The diagram follows possible actions, environment responses, rewards, and successor states whose information contributes back to the starting state.
The Recursive Bellman Relationship
The Bellman equation relates the value of a current state to the expected reward and discounted value of possible successor states. It is recursive because the value on the right-hand side includes the value function applied to states that may be reached later. Those successor-state values are evaluated using the same kind of relationship.
vπ(s) = expected value of [immediate reward + γ × value of the successor state]For a small finite set of possible outcomes, the expectation can be written as a weighted sum. Each outcome contributes its probability multiplied by its immediate reward plus its discounted successor-state value. If the policy can select several actions, the state's value also reflects the actions selected by that policy.
States and State-Action Pairs
A state describes the situation in which the agent is being evaluated. A state-action pair adds the particular action under consideration in that state. This distinction matters in a backup diagram: an open circle identifies a state, while a solid circle identifies a state-action pair. The state connects to the actions available under the policy, and each action connects to its possible rewards and successor states.
A state-level value summarizes what may happen from the state under the policy. A state-action value keeps the chosen action visible, so different actions from the same state can be examined separately.
Rewards, Probabilities, and Discounting
Three ingredients determine the contribution of a possible outcome. The immediate reward describes what is received from the transition. The successor-state value describes what is expected after that transition. The discount factor γ scales the future value, and the probability of the outcome determines how much that outcome contributes to the expectation.
| Ingredient | Role in the calculation |
|---|---|
| Immediate reward | Contributes what is received now |
| Discount factor γ | Scales the successor-state value |
| Successor-state value | Represents the value after the transition |
| Outcome probability | Weights how much a possible outcome contributes |
The four parts of an illustrative Bellman backup
A Small Numerical Backup
Consider one action from state s with two possible outcomes. Outcome 1 has probability 0.6, immediate reward 4, and successor-state value 10. Outcome 2 has probability 0.4, immediate reward 1, and successor-state value 5. Use a discount factor of 0.5. This is a generated illustration of how the expectation can combine multiple possible outcomes; it is not a numerical example taken from the source.
Calculating the Current State Value
Outcome 1 has probability 0.6, reward 4, and successor-state value 10. Outcome 2 has probability 0.4, reward 1, and successor-state value 5. The discount factor is 0.5. Find the expected value of state s.
Calculate outcome 1's backed-up return: Add the immediate reward to the discounted successor value: 4 + 0.5 × 10 = 9.
Weight outcome 1 by its probability: Multiply 9 by 0.6: 0.6 × 9 = 5.4.
Calculate outcome 2's backed-up return: Add the immediate reward to the discounted successor value: 1 + 0.5 × 5 = 3.5.
Weight outcome 2 by its probability: Multiply 3.5 by 0.4: 0.4 × 3.5 = 1.4.
Combine the possible outcomes: Add the weighted contributions: 5.4 + 1.4 = 6.8.
The illustrative value of state s is 6.8.
vπ(s) = 0.6 × (4 + 0.5 × 10) + 0.4 × (1 + 0.5 × 5) = 6.8
Mistakes in Bellman Backups
Treating the current state as independent from future states.
The Bellman relationship connects the current state's value to the expected reward and discounted values of possible successor states.
Fix:
For every possible outcome, include both the immediate reward and the discounted successor-state value.Forgetting probabilities when several outcomes are possible.
The value is an expectation, so possible outcomes contribute according to their probabilities.
Fix:
Multiply each outcome's return by its probability before adding the contributions.Discounting the immediate reward instead of the future value.
The illustrative Bellman structure adds the immediate reward to the discounted successor-state value.
Fix:
Use reward + γ × successor-state value for each outcome.Confusing a state with a state-action pair in a backup diagram.
The source's diagrams distinguish open-circle states from solid-circle state-action pairs.
Fix:
Track whether each node represents the situation itself or a particular action taken in that situation.
Practice the Backup
A state has two possible outcomes under a policy. Outcome 1 has probability 0.7, reward 2, and successor-state value 8. Outcome 2 has probability 0.3, reward 5, and successor-state value 4. Using a discount factor of 0.5, calculate the expected value of the state.
Hints
- First calculate reward + 0.5 × successor-state value for each outcome.
- Multiply each result by its outcome probability.
- Add the two weighted contributions.
What do you think happens?
Before calculating, predict which outcome contributes more to the final expected value.
Reveal answer
Answer: Outcome 1
Outcome 1 contributes 0.7 × (2 + 0.5 × 8) = 4.2, while outcome 2 contributes 0.3 × (5 + 0.5 × 4) = 2.1.
Summary
- The Bellman equation is a recursive consistency relationship for value functions.
- A state's value depends on expected rewards and discounted values of possible successor states.
- Probabilities weight the contributions of different possible outcomes.
- Backup diagrams show how states, state-action pairs, actions, rewards, and successor states are connected.
- The value function vπ is the unique solution to its Bellman equation.
Key Takeaways
- The Bellman equation connects a current state's value to immediate rewards and discounted successor-state values.
- Expected value combines possible outcomes by weighting each outcome according to its probability.
- Backup diagrams make the relationships among states, state-action pairs, actions, rewards, and successor states visible.
- A numerical backup is performed by calculating each outcome's return, weighting it, and adding the contributions.
- Under a policy π, vπ is the unique value function satisfying its Bellman equation.