Gradient Propagation
Data provides the material from which deep-learning systems learn.
Learning Needs Material
A deep-learning system begins with data rather than with intelligence already built in. Data provides the material from which the system can learn. However, having data is only one part of the problem: training methods must also allow a neural network to make useful progress from that material.
The rise of deep learning depended on two improvements happening together. First, machines gained access to much more data. Second, researchers made progress on the algorithms used to train neural networks. Gradient propagation belongs to the second part of this story: it concerns how feedback travels through a network so that the network can improve.
Following the Feedback
A neural network with several layers must use feedback to improve. That feedback has to pass through a stack of layers. Gradient propagation can therefore be understood as following the training signal backward through the network and asking whether it remains strong enough to have a useful effect throughout the stack.
Why More Layers Complicate Training
Adding layers creates a longer path for both the information moving through the network and the feedback used for training. The deeper the stack becomes, the farther the feedback must pass. According to the source, the feedback signal could fade away before it had a useful effect throughout a sufficiently deep network.
Tracing Two Network Depths
Compare the training path in a network with one or two layers of representations with the path in a network containing many layers.
Start with data: In both cases, the system begins with material such as images, videos, or natural-language data.
Pass through representations: The network processes the material through its layers. A shallow network has a shorter stack, while a deep network has many layers.
Send feedback backward: The network must use feedback to improve, and that feedback must pass back through the stack of layers.
Inspect the early layers: In the deeper case, the longer path gives the feedback more opportunity to fade before it reaches earlier layers.
Identify the training limit: If the feedback becomes too weak to have a useful effect throughout the network, training the deep network becomes difficult.
Depth increases the distance that feedback must travel. The central training problem is whether the signal remains useful across that distance.
This explains why neural networks generally remained shallow before a reliable way to train very deep networks was available. They often used only one or two layers of representations and did not consistently outperform more-refined shallow methods such as support vector machines and random forests.
Data Scale and Algorithmic Progress
The internet changed the scale of the material available for machine learning. Internet-scale collection and distribution made very large datasets feasible to collect and distribute. These datasets can include images, videos, and natural-language material, giving learning systems much more material from which to improve.
| Condition | Role in deep learning | Limitation when missing |
|---|---|---|
| Available data | Supplies the material from which a system can learn | The system has less material from which to improve |
| Training algorithms | Determine whether the network can make useful progress from the data | Feedback may not produce useful progress through many layers |
When explaining the rise of deep learning, keep both conditions in view. More data does not by itself solve the problem of training a deep network, and a training method cannot learn from material that is unavailable. The source presents the progress of deep learning as the result of advances in both available data and training algorithms.
Mistakes About Fading Feedback
Assuming that adding layers only adds capacity and cannot affect training.
A deeper stack also creates a longer path for feedback. The source states that feedback could fade across many layers before having a useful effect throughout the network.
Fix:
Treat depth as both a representational opportunity and a training challenge.Treating data and the training algorithm as interchangeable.
Data supplies the material, while training methods determine whether the network can make useful progress from that material.
Fix:
Explain the contribution of data and training methods separately, then describe why both improved together.Thinking that fading feedback means the dataset disappears.
The fading signal is the training feedback moving through the network's layers, not the data supply itself.
Fix:
Keep the two paths distinct: data provides the learning material, while feedback guides improvement through the layer stack.Describing earlier shallow networks as consistently superior to deep networks.
The source says shallow neural networks did not consistently outperform more-refined shallow methods such as support vector machines and random forests before reliable deep-network training was available.
Fix:
Use the qualified claim: shallow neural networks did not consistently outperform those more-refined shallow methods.
Check Your Understanding
A neural network is changed from a shallow stack to a much deeper stack. Explain what happens to the feedback path and why the change can make training difficult. Then explain why access to a large internet-scale dataset does not remove this training challenge by itself.
Hints
- Mention that feedback must pass through the layers.
- Explain what can happen to the signal as the stack becomes deeper.
- Separate the role of data from the role of training algorithms.
What do you think happens?
Before reading the explanation, predict which network creates the greater risk of a feedback signal becoming too weak to guide earlier layers: a network with one or two layers of representations, or a network with many layers.
Reveal answer
Answer: The network with many layers
The feedback must pass through a longer stack in the deeper network, and the source explains that the signal could fade away before having a useful effect throughout that stack.
Key Takeaways
- Data provides the material from which deep-learning systems learn.
- Internet-scale collection and distribution made very large machine-learning datasets feasible, including images, videos, and natural-language material.
- Deep networks are harder to train because feedback must pass through many layers.
- As the layer stack becomes deeper, the feedback signal can fade before it has a useful effect throughout the network.
- The rise of deep learning depended on improvements in both available data and the algorithms used to train neural networks.
Key Takeaways
- Data is the material from which deep-learning systems learn.
- The internet expanded the scale and variety of datasets available for machine learning.
- Adding layers lengthens the path that feedback must travel during training.
- A fading feedback signal can prevent earlier layers of a deep network from receiving a useful training signal.
- Deep learning became more promising when data availability and training algorithms improved together.