REINFORCE Algorithm and Properties
The convergence argument for REINFORCE is based on the expected update, not on every individual episode.
Why One Episode Is Not the Whole Story
REINFORCE learns from sampled episodes. Because different episodes can produce different updates, one episode may suggest a direction that is noisy or temporarily unhelpful. The central convergence idea is therefore about the expected update across the stochastic behavior, not about requiring every individual episode to improve performance.
From Episodes to the Expected Direction
REINFORCE is considered a stochastic gradient method because its update varies from one sampled episode to another. The randomness enters through the sampled episode and its return. Each episode produces an update, but the updates are not identical. When this stochastic behavior is considered in expectation, the expected update has the same direction as the performance gradient. That expected direction is what connects REINFORCE's random episode updates to improvement in expected performance.
The word expected is essential. REINFORCE does not require every sampled episode to produce a helpful update. Its theoretical direction is a statement about the average behavior of the stochastic updates.
A Single Update Versus the Average
Interpreting three sampled episodes
Imagine that three sampled episodes produce three different update directions. How should these individual directions be interpreted?
Episode one: Its update points toward better performance, but it is only one stochastic sample.
Episode two: Its update differs from the first. The difference is evidence of sampling variation, not a failure of the expected-gradient idea.
Episode three: Its update may point in an unhelpful direction. REINFORCE's convergence argument does not require this individual update to improve performance.
Expected behavior: When the stochastic updates are considered in expectation, their direction matches the performance gradient.
A noisy individual episode update and a favorable expected update can exist at the same time.
Step Size and Local Convergence
The expected direction alone is not enough to guarantee the stated improvement result. When the step-size parameter α is sufficiently small, the expected direction of the REINFORCE update supports improvement in expected performance. A correct direction can still fail to provide that guarantee if the step is too large.
The stronger convergence statement uses decreasing α together with standard stochastic approximation conditions. Under those conditions, REINFORCE is assured to converge to a local optimum. This is a local guarantee: it does not claim that REINFORCE must find the best possible policy in every problem.
Why Learning Can Be Slow
REINFORCE is also a Monte Carlo method. Monte Carlo methods use random sampling to approximate solutions, and REINFORCE relies on sampled episodes. The update suggested by one episode can therefore differ substantially from the update suggested by another. This difference is high variance.
High variance makes the path toward the favorable expected direction noisy. Updates may point in noticeably different directions from episode to episode, so useful progress can be slow even when the theoretical convergence result is favorable. This is why it is misleading to label REINFORCE simply stable or unstable: its expected behavior is favorable, while its sampled route can be noisy and slow.
Consider two training runs that both use REINFORCE under conditions supporting convergence. The theoretical statement concerns where the method is expected to go eventually. It does not imply that both runs will show smooth improvement at every episode or reach useful performance quickly. Sampling variation can make the practical route uneven.
| Question | Theoretical convergence | Practical learning speed |
|---|---|---|
| What is being considered? | The expected update and its direction | The sequence of sampled episode updates |
| What does variation mean? | Individual variation does not invalidate the expected direction | High variance can make progress noisy and slow |
| What is the outcome? | Under decreasing α and standard stochastic approximation conditions, convergence to a local optimum | The method may take a long time to show useful progress |
Mistakes in Reading REINFORCE
Assuming every episode must improve performance.
The convergence argument concerns the expected update, not every individual episode.
Fix:
Evaluate the expected behavior of the stochastic updates rather than demanding improvement from each sample.Treating the correct expected direction as a guarantee for any step size.
The improvement statement depends on a sufficiently small step size.
Fix:
Include the step-size condition when stating the expected-performance improvement result.Interpreting convergence as a guarantee of the globally best policy.
The convergence guarantee is to a local optimum.
Fix:
Describe the result as local convergence, not guaranteed global optimality.Calling REINFORCE practically fast because its theory is favorable.
Monte Carlo sampling can create high variance, making learning slow despite favorable theoretical convergence.
Fix:
Discuss theoretical direction and practical variance as two properties that must be considered together.
Check Your Understanding
Explain why the following two statements can both be true: an individual REINFORCE episode produces an unhelpful update, and REINFORCE has an expected update in the direction of the performance gradient.
Hints
- Focus on the meaning of expected update.
- Separate one sampled episode from the average behavior of the stochastic updates.
- Mention how high variance affects the route taken during learning.
State the step-size conditions associated with the two claims: improvement in expected performance and convergence to a local optimum.
Hints
- The first claim uses a sufficiently small α.
- The stronger convergence claim uses decreasing α together with standard stochastic approximation conditions.
What do you think happens?
If two sampled episodes produce substantially different updates, does that by itself contradict the REINFORCE convergence argument?
Reveal answer
Answer: No, because the argument concerns the expected update.
REINFORCE updates can vary from episode to episode. The expected update has the same direction as the performance gradient, while Monte Carlo variation can make individual updates noisy and learning slow.
The REINFORCE Takeaway
- REINFORCE is a stochastic gradient method because its episode-based updates vary across samples.
- The expected update has the same direction as the performance gradient, even though an individual episode update may be noisy or unhelpful.
- A sufficiently small α supports improvement in expected performance.
- Decreasing α together with standard stochastic approximation conditions supports convergence to a local optimum.
- Monte Carlo high variance can make practical learning slow, so theoretical convergence should not be confused with fast or smooth learning.
Key Takeaways
- REINFORCE's convergence argument is about the expected update rather than every individual episode.
- The expected update points in the same direction as the performance gradient.
- Sufficiently small step sizes support expected-performance improvement, while decreasing step sizes and standard stochastic approximation conditions support local convergence.
- Monte Carlo sampling can create high variance, making practical learning noisy and slow.
- Theoretical convergence and practical learning speed are different properties and must be evaluated separately.