Direct Reinforcement Learning
Dyna-Q gives an agent both direct reinforcement learning and planning.
One Real Step, Two Forms of Learning
Direct reinforcement learning uses the agent's actual experience in the environment to improve its behavior. Dyna-Q adds planning to that direct learning. After an agent takes a real step in the maze, it can both learn directly from that experience and perform additional planning updates based on simulated experiences. The parameter n makes the amount of planning explicit.
The central idea is not that direct learning is replaced. Instead, Dyna-Q gives the agent both direct learning from real experience and planning effort after that real step.
Reading the Planning Parameter
In Dyna-Q, n is the number of planning steps performed for each real step through the environment. It therefore measures planning effort per real step. An agent with n = 0 performs no planning and serves as the nonplanning direct-learning baseline. Agents with n = 5 or n = 50 perform five or fifty planning steps, respectively, after each real step.
Comparing Three Planning Settings
Suppose each agent takes one real step in the maze. How much planning does each setting add?
n = 0: The agent performs zero planning steps. This is the nonplanning direct-learning baseline.
n = 5: The agent performs five planning steps after the real step.
n = 50: The agent performs fifty planning steps after the real step.
The settings differ in planning effort per real step: none, five, or fifty planning steps.
Maze Learning Speed
The maze experiment compared how quickly agents improved when n was 0, 5, or 50. Increasing n made performance improve more rapidly. The n = 0 agent was by far the slowest, taking about 25 episodes to reach near-optimal performance. The n = 5 agent needed about five episodes, while the n = 50 agent needed only about three.
| Agent | Planning steps per real step | Episodes to near-optimal performance |
|---|---|---|
| n = 0 | 0 | about 25 |
| n = 5 | 5 | about 5 |
| n = 50 | 50 | about 3 |
Reported learning-speed comparison in the maze experiment
Three Historical Phases
Research on teams of reinforcement learning agents developed through three broad phases. The first phase studied non-associative learning automata in bandit, team, and game problems. The second phase extended learning automata to associative or contextual learning and connected them with artificial neural networks. The third phase was influenced by neuroscience and examined learning rules in relation to synaptic plasticity and other biological constraints.
| Phase | Main focus | What was added |
|---|---|---|
| First | Non-associative learning automata | Learning in bandit, team, and game problems |
| Second | Associative or contextual learning | Context and connections with artificial neural networks |
| Third | Neuroscience-influenced learning | Attention to synaptic plasticity and biological constraints |
The three broad phases of research on teams of reinforcement learning agents
From Automata to Context
Non-associative learning automata formed the focus of the first phase. They were studied in bandit, team, and game problems, but the source distinguishes this work from the contextual bandit case. Associative or contextual reinforcement learning added an explicitly associative context: the learner could be understood in relation to the situation in which action selection occurred.
The second phase also connected associative stochastic learning automata with single-layer artificial neural networks. In this work, the networks received one global reinforcement signal, and the neuron-like learning elements were called associative search elements, or ASEs.
Why A R - P Mattered
Barto and Anandan introduced the associative reward-penalty algorithm, abbreviated A R - P, in 1985. It was important because it was a more sophisticated associative reinforcement learning algorithm that connected several lines of work: stochastic learning automata, pattern classification, associative reinforcement learning, and artificial neural networks.
A R - P extended the associative direction beyond a single learning unit. Teams of A R - P units were connected into multi-layer neural networks. Reported results showed that these teams could learn nonlinear functions, including XOR, while receiving a globally broadcast reinforcement signal. Later, Williams mathematically analyzed and broadened this class of learning rules and showed in 1992 that a special case of A R - P is a REINFORCE algorithm.
Neuroscience Shapes the Third Phase
The third phase was influenced by growing neuroscience support for the possibility that this kind of learning occurs in the brain. Researchers began paying more attention to synaptic plasticity and to constraints suggested by neuroscience, rather than treating a learning rule only as an abstract computational procedure.
- Spike-timing-dependent plasticity, or STDP
- Dopamine
- Reward-modulated STDP
- Synaptic plasticity
- Other biological constraints on learning rules
The important change was a change in the questions researchers asked. Earlier work could evaluate whether an abstract learning rule solved a computational problem. The neuroscience-influenced phase also asked whether the rule was compatible with findings about the brain, including mechanisms related to synaptic plasticity, STDP, dopamine, and reward-modulated STDP.
Check Your Understanding
What do you think happens?
Which agent is the nonplanning direct-learning baseline: n = 0, n = 5, or n = 50?
Reveal answer
Answer: n = 0
The n = 0 agent performs no planning steps. It still learns directly from real experience, so it is a nonplanning direct-learning baseline.
Explain in your own words why the n = 50 agent reached near-optimal performance in about three episodes in the reported maze experiment, while the n = 0 agent took about 25 episodes.
Hints
- Start by defining what n counts.
- Compare the amount of planning performed after each real step.
- Remember that the episode counts came from a particular experiment and its reported conditions.
Place each description in the correct historical phase: non-associative learning automata; associative or contextual learning connected with artificial neural networks; neuroscience-informed attention to synaptic plasticity and biological constraints.
Hints
- The first phase studied bandit, team, and game problems.
- The second phase added context and neural-network connections.
- The third phase was influenced by findings about the brain.
Treating n = 0 as an agent that does not learn.
It still performs direct reinforcement learning from real experience; it simply performs zero planning steps.
Fix:
Describe n = 0 as direct learning without planning.Interpreting n as the number of real steps through the maze.
n controls planning effort per real step.
Fix:
Say that five planning steps are performed after each real step.Treating the maze episode counts as universal results.
The reported comparison came from a particular experiment whose curves averaged 30 repetitions under stated parameter settings.
Fix:
Report the approximate comparison in the context of that experiment.Calling the first phase associative or contextual learning.
The source identifies the absence of the contextual bandit case as a boundary of the non-associative first phase.
Fix:
Reserve associative or contextual learning for the second phase.Reducing A R - P to a single isolated learning unit.
Its historical importance included connecting associative reinforcement learning with artificial neural networks and enabling teams of units to learn nonlinear functions.
Fix:
Explain A R - P as a development that strengthened the connection among automata, associative learning, and neural networks.
Key Takeaways
- Dyna-Q combines direct reinforcement learning from real experience with planning based on simulated experiences.
- The parameter n is the number of planning steps performed per real environment step.
- The n = 0 agent is a nonplanning direct-learning baseline; in the reported maze experiment, n = 5 and n = 50 learned much faster.
- Research on teams of reinforcement learning agents progressed from non-associative learning automata, to associative or contextual learning with neural networks, and then toward neuroscience-informed learning rules.
- The associative reward-penalty algorithm connected stochastic learning automata, associative reinforcement learning, pattern classification, and artificial neural networks.
- The third phase was influenced by neuroscience findings and constraints involving synaptic plasticity, STDP, dopamine, and reward-modulated STDP.
Key Takeaways
- Dyna-Q preserves direct learning while adding planning after each real step.
- n measures planning effort per real step: n = 0 means no planning, while larger values add more planning updates.
- In the reported maze experiment, learning became faster as n increased: about 25 episodes for n = 0, about 5 for n = 5, and about 3 for n = 50.
- The history of team reinforcement learning research moved from non-associative automata to associative neural-network learning and then toward neuroscience-informed rules.
- A R - P was important because it connected associative reinforcement learning with neural networks and later included a special case related to REINFORCE.