Value Function Updates
The reinforcement signal is the number that most directly guides changes in a policy or value function.
The Number That Directs Learning
An agent does not change its policy or value function merely because a situation occurred. A numerical reinforcement signal guides the size and direction of that change. It may be positive, negative, or zero, and it acts as a multiplicative factor in a parameter update.
The simplest reinforcement signal is the reward itself, Rₜ. Temporal-difference learning uses a richer signal: the TD error. This combines the immediate reward with a change in predicted value, so learning can be directed by both what happened immediately and how the value estimate changed.
Breaking Down a TD Error
δₜ = Rₜ + γV(Sₜ) − V(Sₜ₋₁)
The two terms should be understood as contributions to one signal, not as two independent learning signals. If both terms are present, the TD error is a mixture. If one term is zero, the remaining term determines the signal.
Reading the Bellman Equation
qπ(s, a) = reward from (s, a) + discounted successor action value under π
An action value is backed up from what can happen after a particular state-action pair. Taking action a in state s can lead to possible successor states s′ and rewards r. The action value qπ(s, a) therefore depends on the transition reward and on the action values available from the successor state. The policy π determines how the next action a′ is selected at s′.
Checking One State
Bellman verification is performed state by state. To check a proposed state-value equation at one state, substitute that state, list its neighboring outcomes, combine the relevant rewards with the values of the states reached from it, and compare the result with the proposed value. A successful check establishes consistency at that state; it does not establish the equation everywhere.
A local Bellman check
A proposed value function assigns a value to state s. The state has possible neighboring outcomes with known rewards and state values. Verify the equation at s.
Identify the current state: Select the particular state whose proposed value is being checked.
List neighboring outcomes: Record the rewards and successor-state values associated with the outcomes reachable from that state.
Combine the outcomes: Use the Bellman equation to combine the relevant rewards and successor values.
Compare: If the computed result equals the proposed value for s, the equation holds for this state.
The verification is local: it tests the Bellman relationship at the selected state, not automatically at every state.
Shifting Every Reward
Suppose the same constant c is added to every reward. The value of every state then increases by one common amount, written vᶜ. The values shift, but their relative ordering under a policy does not change. A state that was valued above another remains valued above it after the common shift.
Common Reasoning Errors
Treating the reward as the only possible reinforcement signal.
A TD error also contains the secondary contribution γV(Sₜ) − V(Sₜ₋₁).
Fix:
Write the complete TD-error expression before deciding which contributions are present.Calling primary and secondary reinforcement two separate learning signals.
They are conceptual parts of one TD-error reinforcement signal.
Fix:
Add the two contributions to obtain δₜ.Calling a TD error pure primary when the reward is nonzero but the value-difference term is also nonzero.
Both nonzero terms make the signal a mixture.
Fix:
Classify the signal by checking whether each term is zero or nonzero.Assuming one successful Bellman check proves the value function everywhere.
Bellman verification is performed state by state.
Fix:
Repeat the check for every state that must satisfy the equation.Assuming a common reward shift preserves the numerical values.
Every state value increases by a common amount vᶜ.
Fix:
Distinguish the changed absolute values from the unchanged relative ordering.
Apply the Checks
For each situation, identify the reinforcement pattern or value-function consequence. First, Rₜ is zero while γV(Sₜ) − V(Sₜ₋₁) is nonzero. Second, both terms in δₜ are nonzero. Third, every reward is increased by the same constant c. Fourth, a proposed Bellman equation matches at one state but has not been checked elsewhere.
Hints
- A zero reward removes the primary contribution but does not remove a nonzero secondary contribution.
- Two nonzero TD-error terms form a mixture.
- A common reward shift changes every state value by the same amount and preserves relative ordering.
- A check at one state establishes only a local verification.
- A reinforcement signal is the number that most directly guides a policy or value-function update. In TD learning, δₜ combines primary reinforcement Rₜ with secondary reinforcement γV(Sₜ) − V(Sₜ₋₁). Action values are backed up from transition rewards and successor action values, with the policy determining successor actions. Bellman equations must be checked state by state. Adding the same constant to every reward shifts every state value by a common amount while preserving the relative ordering of states.
Key Takeaways
- The reinforcement signal controls the size and direction of a policy or value-function update.
- A TD error combines primary reinforcement from the reward with secondary reinforcement from the temporal-difference value change.
- The action value qπ(s, a) is backed up from the reward and successor action values selected according to policy π.
- Bellman verification is local unless the equation is checked for every state.
- Adding a common constant to all rewards shifts all state values equally without changing their relative ordering.