Advanced deep-learning best practices
Deep learning became more popular through the combined influence of hardware, data, algorithms, investment, and democratization.
Why Deep Learning Accelerated
Deep learning became especially prominent through several reinforcing developments rather than through one isolated breakthrough. Advances in hardware made computation more capable, greater data availability supplied more material for learning, algorithmic improvements strengthened the methods, investment increased activity in the field, and democratization made deep learning more accessible. These influences help explain the recent popularity of deep learning.
A useful first principle is that deep learning's popularity reflects a combination of enabling conditions: better computational hardware, more data, improved algorithms, investment, and wider access.
From Scalars to Tensors
Tensors organize the data used by neural networks. A scalar is a zero-dimensional tensor: it represents one value. A vector is a one-dimensional tensor that can be understood as an ordered collection of values. A matrix is a two-dimensional tensor that organizes values across two dimensions. Higher-dimensional tensors extend this arrangement to additional dimensions. The dimensions and indices describe how the values are arranged, while data batches are another part of neural-network data representation.
| Representation | Dimensions | Arrangement |
|---|---|---|
| Scalar | 0D | One value |
| Vector | 1D | An ordered collection of values |
| Matrix | 2D | Values organized across two dimensions |
| Higher-dimensional tensor | More than 2D | Values organized across additional dimensions |
A conceptual comparison of tensor representations
The important distinction is not merely the number of values. It is the structure used to arrange them. Tensor dimensions and indices describe where values belong within that structure, and tensor operations use those representations as the input to neural-network computation.
Four Ways Tensors Change
Tensor operations are computational steps that manipulate data representations inside neural networks. Element-wise operations calculate on values in corresponding positions. Broadcasting allows an operation involving tensors with different shapes when the operation's rules permit those tensors to work together. Tensor dot performs a tensor calculation that combines values according to the operation. Reshaping changes how the tensor is organized without being described as the same kind of value calculation. These operations belong to the mathematical machinery used by neural networks.
Following One Tensor Through Operations
Track what kind of change occurs when an input tensor passes through four conceptual operations.
Start with organized data: The neural network receives data represented as a tensor, whose dimensions and indices describe how its values are arranged.
Apply an element-wise operation: Values are calculated according to corresponding positions in the tensor.
Use broadcasting: The operation works with tensors of different shapes when the operation's rules allow them to work together.
Apply tensor dot: Values are combined through a tensor dot calculation, producing another tensor result.
Reshape the result: The tensor's organization changes so its values are arranged according to a different shape.
Tensor operations can change values, combine representations, or change organization. They provide the computational path through which neural-network data is transformed.
The Training State Change
Training is a repeated state change. The network begins an iteration with current parameters. Those parameters are used in tensor computations, the result is evaluated with a loss function, gradients are computed, and an optimization method updates the parameters. The updated parameters affect later tensor computations.
Backpropagation is the algorithm used to compute gradients by chaining derivatives through the operations. Gradients provide information for changing model parameters. Stochastic gradient descent is identified as a gradient-based optimization method that uses this gradient information during repeated updates.
Connecting Computation and Optimization
Tensors and optimization answer different questions. Tensor operations describe how data is represented and transformed. Gradient-based optimization describes how the network's parameters change during training. Neither replaces the other: tensor operations produce the computations whose results contribute to the loss, while gradients provide information for changing the parameters used by those computations.
One Complete Learning Cycle
Trace the relationship between an input tensor, tensor operations, loss, gradients, and parameter updates.
Represent the input: Organized data enters the network as a tensor.
Transform the data: Tensor operations manipulate the representation and values as the network performs its computations.
Evaluate the result: A loss function measures the result of the network's current computation.
Compute gradients: Backpropagation chains derivatives through the operations to compute gradients.
Update parameters: A gradient-based optimization method, such as stochastic gradient descent, uses the gradients to update the model parameters.
Begin the next cycle: The updated parameters affect the tensor computations performed in later training iterations.
Learning connects a forward computation with a parameter update: tensors carry and transform organized data, while gradients guide changes to the parameters that control those transformations.
The central trace is: organized data enters as tensors, tensor operations transform it, a loss function evaluates the result, backpropagation computes gradients, and an optimization method updates parameters. Those updated parameters affect later tensor computations.
Mistakes in the Mental Model
Treating every tensor operation as a change to the tensor's shape
The source distinguishes operations that calculate on values from reshaping, which changes organization.
Fix:
Ask whether the operation is calculating on values, allowing differently shaped tensors to work together, combining values through tensor dot, or changing organization through reshaping.Treating tensors and gradients as interchangeable ideas
Tensors represent and transform data, while gradients provide information for changing parameters.
Fix:
Keep the roles separate: tensor operations perform computations, and gradient-based optimization changes parameters.Describing backpropagation as the parameter-update method
Backpropagation computes gradients by chaining derivatives through operations. An optimization method uses those gradients to update parameters.
Fix:
Describe the sequence as backpropagation computes gradients, then stochastic gradient descent or another gradient-based method uses them for updates.Explaining deep learning's popularity through only one cause
The source identifies hardware, data, algorithms, investment, and democratization as contributing factors.
Fix:
Use the combined explanation and describe the factors as reinforcing contributors.
Practice the Learning Trace
A neural network receives organized data, performs tensor operations, evaluates a loss, computes gradients, and changes its parameters. Explain the role of each stage and identify which stage changes the model's parameters.
Hints
- Separate the representation of data from the process of updating parameters.
- Recall what backpropagation computes.
- Identify the optimization method that uses gradients for updates.
Classify each description as scalar, vector, matrix, or higher-dimensional tensor: one value; an ordered collection of values; values arranged across two dimensions; values arranged across additional dimensions. Then explain why dimensions and indices matter when neural-network data is represented.
Hints
- Start with the number of dimensions.
- A vector is one-dimensional and a matrix is two-dimensional.
- Higher-dimensional tensors extend the arrangement beyond two dimensions.
Key Takeaways
- Deep learning's recent popularity came from the combined influence of hardware, data, algorithms, investment, and democratization.
- Tensors organize neural-network data from zero-dimensional scalars through vectors, matrices, and higher-dimensional structures.
- Element-wise operations, broadcasting, tensor dot, and reshaping manipulate tensor values or organization in different ways.
- Backpropagation computes gradients by chaining derivatives through operations, while stochastic gradient descent uses gradient information to update parameters.
- Training repeatedly connects tensor computation, loss evaluation, gradient computation, and parameter updates.
Key Takeaways
- Deep learning became more prominent through the combined effects of hardware, data, algorithms, investment, and democratization.
- Tensors provide the structured data representations used by neural networks.
- Tensor operations transform representations and values, while reshaping changes organization.
- Backpropagation computes gradients and stochastic gradient descent uses them to update model parameters.
- The training process is a repeated loop in which updated parameters affect later tensor computations.