Why Deep Learning Now?
Between 1990 and 2010, off-the-shelf CPUs became approximately 5,000 times faster, making small deep-learning models runnable on laptops.
The Hardware Turning Point
The rise of deep learning was not caused only by better algorithms. Hardware changed the boundary between an idea that could be computed and an idea that was impractical to run. Between 1990 and 2010, off-the-shelf CPUs became approximately 5,000 times faster. That improvement made small deep-learning models runnable on an ordinary laptop, even though the same kind of task would have been intractable 25 years earlier.
From Runnable to Impractical
The CPU improvement did not make every deep-learning workload easy to run on a laptop. It made small deep-learning models possible on ordinary hardware. Typical computer-vision and speech-recognition models still require far more computational power than a laptop can provide. This creates an important distinction: a model can be small enough to run locally, while a typical model in an application area can still exceed the available computing capacity.
Reading the Hardware Story
Use the source facts to explain why the phrase deep learning now refers to both progress and limitation.
Start with CPU progress: Between 1990 and 2010, off-the-shelf CPUs became approximately 5,000 times faster.
Identify the new possibility: That improvement made small deep-learning models runnable on an ordinary laptop.
Identify the remaining limitation: Typical computer-vision and speech-recognition models still require far more computational power than a laptop can provide.
Look for a different kind of capacity: Because the remaining workloads contain many parallelizable calculations, GPUs became an important match for them.
CPU progress made small models practical, while GPU-based parallel processing addressed workloads that remained too demanding for a laptop.
Why Neural Networks Fit Parallel Hardware
A deep neural network is well suited to a parallel processor because it consists mostly of many small matrix multiplications. The individual computational pieces can be handled in parallel instead of being treated as one long, purely sequential calculation. This property is called parallelizability. It connects the structure of the model with the design goal of a GPU: performing many operations at the same time.
From Graphics to General Computation
GPUs were not originally created for deep learning. They were developed to render video-game graphics, including increasingly photorealistic scenes. That investment produced fast, massively parallel chips. Although their original purpose was graphics, the hardware was also useful for other applications whose computations could be parallelized.
The transition was therefore a change in use, not a claim that GPUs were originally built for neural networks. Their graphics role encouraged the development of massively parallel hardware. Once researchers recognized that other workloads also contained computations that could be parallelized, the same kind of hardware became useful beyond graphics.
CPU and GPU Work Styles
The useful contrast is between a CPU's general-purpose role and a GPU's ability to perform many operations at the same time. Deep-learning computation benefits from the GPU style because the model contains many similar matrix-multiplication pieces. The GPU supplies a different kind of computational capacity rather than simply being described as a universally better replacement for a CPU.
CUDA Makes the Hardware Usable
The hardware alone was not the whole story. In 2007, NVIDIA launched CUDA, a programming interface for its GPUs. CUDA enabled researchers and developers to use those GPUs for applications beyond their original graphics role. The source describes an early consequence in scientific computing: a small number of GPUs began replacing massive clusters of CPUs in highly parallelizable applications, starting with physics modeling. Because deep neural networks are also highly parallelizable, researchers began writing CUDA implementations of neural networks around 2011.
The Four-Part Fit
Trace the sequence that connected ordinary laptops, GPUs, CUDA, and deep learning.
CPU improvement: Faster off-the-shelf CPUs expanded what an ordinary laptop could attempt and made small deep-learning models runnable.
GPU development: Graphics rendering drove the development of fast, massively parallel chips that could also support other parallelizable workloads.
Software access: CUDA, launched by NVIDIA in 2007, gave researchers and developers a programming interface for using NVIDIA GPUs beyond graphics.
Model fit: Deep neural networks contain many small matrix multiplications, so their computation is highly parallelizable and matches the GPU design goal of performing many operations at once.
The rise of deep learning depended on a sequence of fit: faster CPUs expanded laptop capability, GPUs supplied parallel capacity, CUDA exposed that capacity to software, and neural-network structure made the capacity useful.
Mistakes in the Hardware Story
Assuming CPU improvement made typical deep-learning models easy to run on a laptop.
Typical computer-vision and speech-recognition models still require far more computational power than a laptop can provide.
Fix:
Distinguish between small models that became practical and typical models that remained too computationally demanding.Saying that GPUs were originally created for deep learning.
Their later usefulness for deep learning came from the fact that their massively parallel design also suited highly parallelizable workloads.
Fix:
Describe deep learning as a later use of graphics-oriented parallel hardware.Treating CUDA as the GPU itself.
CUDA is software access to the hardware, not the hardware component.
Fix:
Explain the relationship as GPU hardware plus CUDA programming support.Ignoring the model structure when explaining GPU usefulness.
Without the parallelizable structure of those calculations, the connection between neural networks and GPUs is incomplete.
Fix:
Always connect GPU usefulness to the many computational pieces that can be handled in parallel.
Check the Causal Chain
Explain in four linked steps why deep learning became more practical during this period. Your answer should include CPU improvement, GPU development, CUDA, and the parallel structure of deep neural networks.
Hints
- Begin with what approximately 5,000-fold CPU improvement changed on an ordinary laptop.
- Then distinguish small models from typical computer-vision and speech-recognition models.
- Explain what GPUs were originally designed to do and why their hardware suited other parallelizable workloads.
- Finish by identifying CUDA as a programming interface and matrix multiplications as the model structure that fits GPU processing.
The Complete Explanation
- Between 1990 and 2010, off-the-shelf CPUs became approximately 5,000 times faster, making small deep-learning models runnable on ordinary laptops.
- Typical computer-vision and speech-recognition models still require far more computational power than a laptop can provide.
- GPUs began as video-game graphics hardware, but their massively parallel design also suited scientific and deep-learning workloads.
- CUDA, launched by NVIDIA in 2007, provided a programming interface that enabled applications to use NVIDIA GPUs beyond graphics.
- Deep neural networks match GPUs because they consist mostly of many small matrix multiplications whose computational pieces can be handled in parallel.
Key Takeaways
- Hardware progress helped move deep learning from impractical to runnable for some workloads.
- CPU improvement made small models possible on laptops, but typical vision and speech models still exceeded laptop capacity.
- GPUs supplied massively parallel computation after beginning as graphics-rendering hardware.
- CUDA made NVIDIA GPU capacity available to software beyond graphics.
- The many matrix multiplications in deep neural networks created a strong match with GPU parallel processing.