Differentiable Value Functions
Gradient Monte Carlo uses gradient descent to update the weights of a value function.
Learning from Complete Episodes
A value function estimates how much return can be expected from a state. Gradient Monte Carlo learns such an estimate by representing the value function with weights and improving those weights from experience. The policy π being evaluated generates an episode, and the observed return G_t provides a target for the estimate at each visited state.
The value estimate is written as v̂(S_t, θ), where S_t is the state visited at time t and θ is the current set of weights. For every time step from 0 through T − 1, Gradient Monte Carlo compares the observed quantity G_t with v̂(S_t, θ). It then adjusts θ. Because the adjustment uses the gradient ∇v̂(S_t, θ), the value function must be differentiable with respect to its weights.
Episode-to-Update Flow
- Use the policy π being evaluated to generate an episode.
- Visit each time step t from 0 through T − 1 in that episode.
- Use the episode to obtain G_t for the visited state S_t.
- Evaluate the current estimate v̂(S_t, θ).
- Use the gradient ∇v̂(S_t, θ) together with the difference between G_t and the estimate to adjust θ.
Tracing One Visited State
Suppose an episode generated by π visits state S_t. Explain what information is used before the weights are changed.
Observe the episode: The episode supplies the observed return G_t associated with the visit to S_t.
Estimate the value: The current weights produce the estimate v̂(S_t, θ).
Measure the mismatch: The algorithm considers the difference between G_t and v̂(S_t, θ).
Use sensitivity: The gradient ∇v̂(S_t, θ) identifies how the estimate responds to the weights, so it helps direct the adjustment to θ.
The update uses the observed return, the current estimate, and the value function's gradient.
How the Weights Move
The central learning signal is the difference between G_t and v̂(S_t, θ). If the observed return is different from the current estimate, the weights must be adjusted so that the differentiable value function can better represent the observed return. The gradient ∇v̂(S_t, θ) matters because it describes how the estimated value changes with respect to the weights.
Generated example: If G_t is higher than v̂(S_t, θ), the current estimate is below the observed return. The difference signals that the estimate needs to move upward for this training case. The gradient identifies how changing θ changes the estimate, allowing the update to be directed through the weights rather than by replacing the estimate directly.
Returns Relative to Average Reward
A differential return is a return formed from differences between rewards and the true average reward. Instead of treating a reward by itself as the reference, the continuing-task setting asks whether that reward is above or below the long-run average.
Interpreting Relative Rewards
Use a generated sequence to interpret rewards relative to a true average reward of 5.
Compare the first reward: A reward of 7 is 2 above the average reward of 5.
Compare the next reward: A reward of 3 is 2 below the average reward of 5.
Accumulate the differences: The differential return combines these relative reward differences, rather than treating 7 and 3 as isolated absolute quantities.
The example's differential return reflects performance relative to the average-reward reference.
State and State-Action Values
Differential value functions preserve the conditional-expectation structure of ordinary value functions, but the return being expected is now a differential return. The state value vπ(s) is the expected differential return conditioned on state s while following policy π. The state-action value qπ(s, a) is the expected differential return conditioned on state s and action a while following policy π.
| Function | Conditioning information | Expected quantity |
|---|---|---|
| vπ(s) | State s | Differential return while following π |
| qπ(s, a) | State s and action a | Differential return while following π |
Bellman Changes for Differential Values
Differential value functions still have Bellman equations. The relationship continues to connect the value now with values at later states or state-action pairs. However, forming the differential versions requires two structural changes to the ordinary Bellman equations.
- Remove the discount factors γ.
- Replace each reward with its difference from the true average reward.
- Keep the continuation value, so the equation still relates the current value to later values.
Mistakes in the Update and Return
Treating v̂(S_t, θ) as the observed return.
Gradient Monte Carlo compares the observed quantity G_t with the current estimate. They play different roles.
Fix:
Keep G_t as the episode-based quantity and v̂(S_t, θ) as the current prediction.Ignoring the gradient.
The method requires the gradient of the differentiable value function with respect to the weights.
Fix:
Use ∇v̂(S_t, θ) to determine how the weights affect the estimate.Confusing a state value with a state-action value.
qπ(s, a) is conditioned on both a state and a selected action, whereas vπ(s) is conditioned on the state.
Fix:
Ask whether the action is fixed: if only the state is fixed, use vπ(s); if the state and action are fixed, use qπ(s, a).Keeping discount factors when forming a differential Bellman equation.
One of the two structural changes is removing the discount factors.
Fix:
Remove γ and replace each reward with its difference from the true average reward.Using an absolute reward instead of a reward difference in a continuing task.
A differential return is built from rewards relative to the true average reward.
Fix:
First compare each reward with the average-reward reference, then accumulate the differences.
Check Your Understanding
An episode is generated using policy π. At one visited state S_t, the observed return G_t differs from v̂(S_t, θ). Explain the role of each of these three items in the update: G_t, v̂(S_t, θ), and ∇v̂(S_t, θ). Then state the two changes needed to form the differential Bellman equations.
Hints
- G_t comes from the episode, while v̂(S_t, θ) comes from the current weights.
- The difference between those two quantities supplies the learning signal.
- The differential Bellman changes concern discount factors and rewards relative to the true average reward.
What do you think happens?
Before checking the explanation, predict whether qπ(s, a) conditions on more information than vπ(s).
Reveal answer
Answer: Yes, because qπ(s, a) conditions on both state and action.
vπ(s) is the expected differential return conditioned on state s. qπ(s, a) is the expected differential return conditioned on state s and action a.
Key Takeaways
- Gradient Monte Carlo generates an episode with policy π and updates the weights for every time step from 0 through T − 1.
- The update uses the observed return G_t, the estimate v̂(S_t, θ), and the gradient ∇v̂(S_t, θ).
- Differential returns accumulate reward differences relative to the true average reward.
- vπ(s) conditions on a state, while qπ(s, a) conditions on a state and an action; both estimate expected differential returns.
- Differential Bellman equations remove discount factors and replace rewards with differences from the true average reward, while retaining the continuation value.
Key Takeaways
- Gradient Monte Carlo learns weights for a differentiable value function by comparing episode returns with current value estimates.
- The gradient of the value function determines how the weights can change the estimate.
- Differential returns measure rewards relative to the true average reward in continuing tasks.
- Differential state and state-action values differ in whether the expectation is conditioned on only a state or on a state-action pair.
- Differential Bellman equations remove discount factors and use reward differences from the average reward.