Evolution of Reinforcement Learning
Reinforcement learning grew from two independently pursued historical directions.
Two Paths to One Field
Reinforcement learning did not begin as one unified subject. It grew from two largely independent historical directions. One direction emphasized learning through trial and error. The other focused on optimal control, using value functions and dynamic programming to address control problems, largely without learning. Their later connection helped form the modern field of reinforcement learning.
The historical sequence is best understood as convergence, not as a single invention: trial-and-error learning, optimal control, and temporal-difference methods eventually came together in the late 1980s.
Trial and Error
The trial-and-error thread emphasizes learning through interaction. A system tries behaviors and learns from the results of those trials. The system is therefore understood through the process of trying and learning, rather than through a purely analytical solution to a control problem.
| Viewpoint | Central emphasis | How behavior is understood |
|---|---|---|
| Trial-and-error learning | Learning from trials | The system improves through interaction and the results of experience |
| Optimal control | Solving a control problem | The problem is addressed through value functions and dynamic programming |
Value Functions and Backups
The optimal-control thread centered on value functions and dynamic programming. A value function represents the value associated with a state or action. Dynamic programming uses these value-oriented ideas to work through a control problem. Together, they provide a route toward selecting control behavior without making learning from trial and error the central mechanism.
Do not treat the optimal-control thread as identical to trial-and-error learning. The source describes optimal control as developing largely without learning, while the trial-and-error thread places learning at the center.
The Temporal-Difference Bridge
The two main threads were not completely separate. Temporal-difference methods formed a less sharply defined third thread and helped connect trial-and-error learning with value-oriented ideas from optimal control. Their historical importance is therefore not that they replace either thread, but that they provide a bridge between learning from experience and improving value estimates.
Reading the Historical Bridge
Suppose a learner encounters a method that uses experience from interaction while also improving estimates associated with states or actions. Which historical threads does this method connect?
Identify the learning source: Because the method uses experience from interaction, it reflects the trial-and-error thread.
Identify the estimated object: Because the method improves value estimates, it also relates to the value-oriented optimal-control thread.
Name the bridge: Temporal-difference methods are the historical connection between these two emphases.
The method belongs to the connecting thread: it links learning from experience with value-oriented ideas.
Evolutionary Policy Search
Evolutionary methods take a different route from methods that estimate how valuable individual states or actions are. They search over complete policies. Each candidate agent follows its policy while interacting with the environment, receives reward for its lifetime of behavior, and is judged by the result. Candidates associated with better reward are selected as the basis for further search.
Comparing Three Complete Policies
Imagine three non-learning agents, each using a different policy. After the same interaction period, one obtains more reward than the other two.
Run the policies: Each agent follows its own policy while interacting with the environment. The agents do not improve through learning during their individual lifetimes.
Evaluate complete behavior: The method evaluates the reward produced by each agent's complete behavior over its interaction period.
Compare candidates: The candidate associated with more reward is favored over weaker candidates.
Continue the search: The stronger candidate becomes part of the basis for further policy search.
The method selects according to lifetime reward, not according to a separate value estimate for every state or action.
Two Kinds of Evaluation
A value function estimates the value of a state or an action. Evolutionary evaluation asks a different question: how well did an entire policy perform over its interaction period? The unit being evaluated is therefore different. One approach attaches an estimate to a particular state or action; the other scores the complete behavior produced by a policy.
Assuming every reinforcement learning method must construct a value function.
Evolutionary methods can compare complete policies by the reward produced during their lifetimes without constructing a value function.
Fix:
First identify the unit of evaluation: a state or action for value estimation, or an entire policy for evolutionary evaluation.Calling lifetime reward an estimate of one state or action.
The evolutionary method evaluates the complete behavior produced by the policy over its interaction period.
Fix:
Describe the result as a policy-level or lifetime-behavior evaluation.Treating evolutionary candidates as learning during their own lifetimes.
The source describes the agents as non-learning during their individual lifetimes; selection occurs across candidates after behavior is evaluated.
Fix:
Separate individual behavior from the broader search: an agent follows its policy, then candidates are compared and stronger ones are favored.
When Evolutionary Methods Fit
Evolutionary methods may be effective when the policy space is small, when good policies are common or easy to find within that space, or when substantial time is available for searching. These conditions make it more plausible that comparison and selection will encounter useful policies.
They can also be useful when the learning agent cannot accurately sense the state of its environment. Methods that depend on accurate environmental sensing may face difficulty in that situation, while evolutionary evaluation can still compare the behavior produced by different policies according to the reward they obtain.
A problem has a small policy space, limited environmental sensing, and enough time to compare many candidate policies. Explain why evolutionary methods could be considered, and state what they would evaluate.
Hints
- Look for the conditions that make comparison and search more plausible.
- Distinguish a complete policy from a value estimate attached to one state or action.
Historical Convergence
- Reinforcement learning grew from two largely independent threads: trial-and-error learning and optimal control.
- Trial-and-error learning emphasizes interaction and learning from the results of trials.
- Optimal control developed around value functions and dynamic programming, largely without learning.
- Temporal-difference methods helped connect experience-based learning with value-oriented ideas.
- Evolutionary methods evaluate complete policies by lifetime reward and select stronger candidates without requiring state-value or action-value estimation.
Key Takeaways
- Modern reinforcement learning emerged from the convergence of trial-and-error learning, optimal control, and temporal-difference methods.
- Trial-and-error learning focuses on learning through interaction, whereas optimal control focuses on value functions and dynamic programming.
- Temporal-difference methods provided a historical bridge between experience-based learning and value-oriented control.
- Evolutionary methods compare complete policies by the reward produced over their lifetimes instead of estimating the value of individual states or actions.
- Evolutionary approaches may fit small or searchable policy spaces, long search periods, and situations with limited environmental sensing.