Deep Convolutional Neural Networks
DQN combines Q-learning and a deep convolutional ANN in one reinforcement-learning agent.
From Decisions to Visual Input
A reinforcement-learning agent must solve two connected problems. First, it needs a reinforcement-learning method to guide its decisions. Second, it must work with the information supplied by the task. Q-learning addresses the decision-making side, while a deep convolutional artificial neural network helps process spatial input such as an image. A deep Q-network, or DQN, brings these capabilities together in one agent.
The DQN Combination
Q-learning by itself names the reinforcement-learning method used to guide decisions. A DQN is more specific: it combines Q-learning with a deep convolutional artificial neural network. The network serves as the function-approximation component, allowing the agent to work with supplied information such as spatial image input instead of depending only on a table of separately listed values.
| System | Decision method | Input-processing component | Main distinction |
|---|---|---|---|
| Q-learning alone | Q-learning | Not specified by the DQN definition | Names the reinforcement-learning method |
| Deep Q-network | Q-learning | Deep convolutional artificial neural network | Combines decision guidance with a neural network suited to spatial input |
Separating the Two Roles
An agent receives an image-like observation and must choose an action. Which part of a DQN handles the decision-learning method, and which part helps process the supplied spatial information?
Identify the decision method: Q-learning provides the reinforcement-learning method that guides the agent's decisions.
Identify the input-processing component: The deep convolutional artificial neural network processes the spatial input and learns internal representations relevant to the task.
Combine the roles: Together, these components form a DQN rather than Q-learning alone.
A DQN is Q-learning paired with a deep convolutional artificial neural network.
Why Depth Helps Approximation
In reinforcement learning, a neural network can be used to approximate functions involved in the task. Earlier reinforcement-learning successes showed that multilayer neural networks could approximate functions and learn internal representations through backpropagation. Their layers can therefore transform supplied information into representations that are useful for the task, rather than requiring every useful feature to be designed manually in advance.
Depth is useful here because multilayer networks can learn internal, task-relevant representations. The source describes this capability at the level of function approximation and representation learning; it does not require treating the network as a list of specific layer operations.
Spatial Structure and Convolution
A deep convolutional artificial neural network is a many-layered neural network specialized for spatial arrays of data, including images. Spatial data has organization: the arrangement of its values carries information. This makes convolutional structure suitable for the input-processing role in a DQN when the task supplies image-like observations.
Learned and Handcrafted Features
A handcrafted feature is an input feature designed by a person in advance for a particular problem. A learned feature is an internal representation discovered by the network during training because it is useful for the task. Earlier neural-network reinforcement-learning successes could discover task-relevant features, but the strongest demonstrations often relied on specialized handcrafted features. DQNs are important in this lesson because their deep network can learn task-relevant representations instead of depending entirely on manually designed features.
Generated example: Imagine two systems receiving the same spatial input. In one system, a person first decides which input properties should be supplied as features. In the other, a deep multilayer network learns internal representations through training. The first system depends on the person's feature-design choices; the second can discover features relevant to the task through the network's learning process.
Tracing a DQN System
A Complete Conceptual Trace
Trace the roles of the components when a DQN receives a spatial observation and must guide an action.
Receive the observation: The task supplies information represented as a spatial array, such as an image.
Process the spatial input: The deep convolutional artificial neural network is the component suited to processing this kind of organized input.
Learn internal representations: The multilayer network can learn task-relevant internal representations rather than relying entirely on handcrafted features.
Estimate action values: The network is used with Q-learning to provide estimates associated with possible actions.
Guide the decision: Q-learning supplies the reinforcement-learning role that guides the agent's decisions.
The DQN joins spatial-input processing, learned representations, function approximation, and Q-learning in one reinforcement-learning agent.
Common Conceptual Mistakes
Treating Q-learning alone as a DQN
A DQN is defined by combining Q-learning with a deep convolutional artificial neural network.
Fix:
Check for both components: the Q-learning method and the deep convolutional neural network.Assuming depth automatically means image processing
The source distinguishes multilayer networks for function approximation from convolutional ANNs specialized for spatial arrays such as images.
Fix:
Use convolutional terminology when the network's specialization for spatial data is relevant.Assuming all useful features must be manually designed
Deep multilayer networks can learn internal, task-relevant representations.
Fix:
Distinguish learned representations from handcrafted features.Treating TD-Gammon and DQN as the same system
TD-Gammon is presented as an earlier reinforcement-learning and neural-network example, while DQN has the specific Q-learning plus deep convolutional ANN combination.
Fix:
Use each example for its stated historical or conceptual role.
Check Your Understanding
A reinforcement-learning agent receives image observations. Explain why a deep convolutional neural network may be suitable as part of the agent, and explain what Q-learning contributes. Then state one difference between a learned feature and a handcrafted feature.
Hints
- Start with the fact that an image is a spatial array.
- Separate the input-processing role from the reinforcement-learning decision role.
- For the feature comparison, focus on who or what determines the feature before learning.
What do you think happens?
An agent uses Q-learning but has no deep convolutional artificial neural network. Based on the definition used in this lesson, is it already a DQN?
Reveal answer
Answer: No, because a DQN combines Q-learning with a deep convolutional artificial neural network.
Q-learning supplies the reinforcement-learning method, but the DQN definition requires the paired deep convolutional neural-network component as well.
Key Takeaways
- A DQN combines Q-learning with a deep convolutional artificial neural network.
- Multilayer neural networks can approximate functions and learn internal, task-relevant representations through training.
- Convolutional ANNs are suited to spatial arrays such as images because spatial organization carries information.
- Learned features are discovered through training for the task, while handcrafted features are designed by a person in advance.
- TD-Gammon and DQN should not be treated as the same system: the DQN definition specifically names Q-learning and a deep convolutional ANN.
Key Takeaways
- A deep Q-network is more than Q-learning: it combines Q-learning with a deep convolutional artificial neural network.
- Multilayer networks support function approximation and can learn internal representations relevant to the task.
- Convolutional structure is appropriate for spatial arrays such as images because the arrangement of values carries information.
- Deep networks can reduce dependence on features handcrafted for a particular problem by learning task-relevant features during training.