REINFORCE: A Monte Carlo Policy Gradient Method
The convergence argument for REINFORCE is based on the expected update, not on every individual episode.
The Learning Signal
REINFORCE learns by using sampled episodes to adjust a policy. The update produced by one episode is not guaranteed to improve performance. Instead, REINFORCE is understood through the behavior of its updates in expectation: when the stochastic updates are considered together, their expected direction matches the direction of the performance gradient. This distinction explains how REINFORCE can have favorable convergence properties while still learning slowly in practice.
The central idea is expected direction, not guaranteed improvement after every individual episode.
From Episode to Update
A stochastic gradient method uses updates that vary from one episode to another. REINFORCE has this property because it relies on sampled episodes. Each episode supplies information for an update, but the update depends on random sampling and can therefore differ substantially from the update suggested by another episode. The individual update is consequently a noisy estimate of the direction that would be obtained from the full expected behavior of the method.
Two Episodes, Two Different Directions
Illustrate why a single REINFORCE episode should be treated as a noisy estimate rather than as the performance gradient itself.
Episode A: One sampled episode produces an update in one direction.
Episode B: Another sampled episode can produce a substantially different update because Monte Carlo sampling introduces variation.
Consider the updates together: The convergence argument considers the expected update across the stochastic behavior, rather than requiring both individual episodes to point in the same direction.
An individual episode can be noisy even though the expected REINFORCE update has the same direction as the performance gradient.
Why the Average Direction Matters
The expected REINFORCE update has the same direction as the performance gradient. This is the property that makes REINFORCE a stochastic gradient method. The method does not need every sampled update to be helpful. Its theoretical justification comes from the direction obtained when the random updates are considered in expectation. In that expected sense, the method points toward improving performance.
Conditions for Progress
The favorable expected direction is not by itself a guarantee that every step improves expected performance. When the step-size parameter α is sufficiently small, the expected direction of the REINFORCE update supports improvement in expected performance. The size condition matters because a correct direction can still be paired with a step that is too large for the stated improvement guarantee.
The stronger convergence statement requires decreasing α together with standard stochastic approximation conditions. Under those conditions, REINFORCE is expected to converge to a local optimum. This is a local guarantee: it does not claim that REINFORCE must find the best possible policy in every problem.
| Claim | Required condition | Scope |
|---|---|---|
| Improvement in expected performance | α is sufficiently small | Expected behavior |
| Convergence | Decreasing α and standard stochastic approximation conditions | Convergence to a local optimum |
The step-size condition and the convergence claim should not be treated as the same statement.
Variance and Learning Speed
REINFORCE is also a Monte Carlo method, and Monte Carlo methods use random sampling to approximate solutions. Because REINFORCE relies on sampled episodes, the updates can have high variance. The update suggested by one episode may differ substantially from the update suggested by another. These variations can make the route toward the favorable expected direction noisy and slow.
Theoretical Direction, Practical Delay
Explain how high variance can coexist with a favorable expected update.
Observe sampled episodes: The episodes do not all produce the same return or the same update. Their suggested changes can differ substantially.
Combine the stochastic behavior: The expected update still has the same direction as the performance gradient.
Consider the learning path: Because individual updates vary, successive updates can be noisy and may not consistently move in the expected direction.
REINFORCE can have good theoretical convergence properties while requiring many sampled episodes to make practical progress.
Theory versus Practice
| Theoretical convergence | Practical learning speed |
|---|---|
| Concerns the expected update and its direction. | Concerns the noisy path created by sampled episodes. |
| Can support improvement when α is sufficiently small. | Can be slow when Monte Carlo updates have high variance. |
| With decreasing α and standard stochastic approximation conditions, supports convergence to a local optimum. | May require substantial experience before progress becomes evident. |
| Provides a local guarantee rather than a guarantee of the best possible policy. | Describes how quickly useful progress is observed in practice. |
A convergence theorem describes what happens under stated conditions in the relevant expected or asymptotic sense. It does not promise that each episode improves performance or that learning will be fast.
Check Your Understanding
A learner observes that several consecutive episodes produce updates that point in different directions. Should this observation alone be taken as evidence that REINFORCE has no performance-gradient property? Explain the difference between the individual updates and the expected update.
Hints
- Recall what stochastic means in a stochastic gradient method.
- Focus on the word expected.
- Separate the direction of one episode-specific update from the direction of the expected update.
State the two step-size situations discussed for REINFORCE: one that supports improvement in expected performance and the stronger set of conditions associated with convergence to a local optimum.
Hints
- The first statement concerns α being sufficiently small.
- The stronger statement uses decreasing α and standard stochastic approximation conditions.
Assuming every episode-specific update must improve performance.
The convergence argument concerns the expected update, not every individual episode.
Fix:
Evaluate the expected direction of the stochastic updates rather than demanding that each sampled update be helpful.Treating a favorable convergence result as evidence that learning must be fast.
Monte Carlo variance can make the learning path noisy and slow.
Fix:
Keep theoretical convergence and practical learning speed as separate ideas.Describing the guarantee as convergence to the globally best policy.
The stated guarantee is local, not a claim that the best possible policy is always found.
Fix:
Describe the result as convergence to a local optimum.Ignoring the size of α when discussing improvement.
The improvement statement requires α to be sufficiently small.
Fix:
Include the step-size condition when stating the expected-performance improvement claim.
Key Takeaways
- REINFORCE is a stochastic gradient method because its episode-specific updates vary, while their expected update has the same direction as the performance gradient.
- The convergence argument applies to expected behavior, not to every individual episode.
- A sufficiently small α supports improvement in expected performance.
- Decreasing α together with standard stochastic approximation conditions supports convergence to a local optimum.
- Monte Carlo sampling can create high-variance updates, so theoretical convergence and practical learning speed must be distinguished.
Key Takeaways
- REINFORCE estimates a policy-gradient direction from sampled episodes, so its individual updates are stochastic.
- The expected update has the same direction as the performance gradient.
- A sufficiently small step size supports expected-performance improvement, while decreasing step sizes and standard stochastic approximation conditions support convergence to a local optimum.
- High Monte Carlo variance can make progress noisy and slow even when the theoretical convergence behavior is favorable.
- Convergence to a local optimum is not the same as fast learning or finding the globally best policy.