Fundamentals of machine learning
Deep learning became more popular through the combined influence of hardware, data, algorithms, investment, and democratization.
Why Deep Learning Became Prominent
Deep learning became especially prominent through the combined influence of several factors rather than through one isolated breakthrough. Advances in hardware, greater data availability, improvements in algorithms, a new wave of investment, and the democratization of deep learning reinforced one another. Together, these factors helped make deep learning more popular in recent years.
The important idea is combination. The source does not describe popularity as the result of hardware, data, algorithms, investment, or democratization acting alone. It identifies their combined influence.
From Scalars to Tensors
A tensor is a way to organize the data used by a neural network. The organization is described by dimensions and indices. A scalar is a zero-dimensional tensor containing one value. A vector is a one-dimensional tensor, understood as an ordered collection of values. A matrix is a two-dimensional tensor that organizes values across two dimensions. Higher-dimensional tensors extend this arrangement to additional dimensions. Data batches are also part of neural-network data representation.
| Representation | Dimensions | Organization |
|---|---|---|
| Scalar | 0D | One value |
| Vector | 1D | An ordered collection of values |
| Matrix | 2D | Values arranged across two dimensions |
| Higher-dimensional tensor | More than 2D | Values arranged across additional dimensions |
A tensor's dimensions describe how its values are organized.
Operations on Tensor Representations
Tensor operations are the computational steps that manipulate both data representations and values inside a neural network. Four important categories are element-wise operations, broadcasting, tensor dot, and reshaping. They do not all change a tensor in the same way: reshaping changes how values are organized, while element-wise operations and tensor dot perform calculations on values. Broadcasting makes some operations possible when tensors have different shapes and the operation's rules allow them to work together.
Following One Tensor Transformation
A neural network receives organized data and must transform it before producing a result. Trace what each operation contributes without assuming that every operation has the same effect.
Represent: The data enters as a tensor whose dimensions and indices describe how its values are arranged.
Calculate: An element-wise operation can calculate on corresponding values. A tensor dot can calculate across tensor inputs. These are value-transforming operations.
Align: Broadcasting addresses an operation involving tensors with different shapes when the operation's rules allow them to work together.
Reorganize: Reshaping changes the tensor's organization so that the same data representation is arranged differently for a later computation.
Tensor operations are best understood by asking two questions: did the operation calculate new values, or did it change how existing values are organized?
Assuming reshaping is the same as calculating new values.
The source distinguishes reshaping from operations that calculate on values. Reshaping changes the organization of the tensor.
Fix:
Describe reshaping as a change in arrangement, and describe element-wise operations or tensor dot as value calculations.Assuming broadcasting automatically makes any different-shaped tensors compatible.
Broadcasting applies when the operation's rules allow tensors with different shapes to work together.
Fix:
Check whether the operation's rules allow the shapes to work together before describing the operation as valid.Treating all tensor operations as interchangeable.
These categories manipulate representations and values in different ways.
Fix:
Name the operation category and state whether it calculates on values, aligns shapes, or reorganizes the representation.
The Forward Learning State
A neural network does not learn through a single operation. At one point in the training cycle, the network uses its current parameters to perform tensor computations. Organized input data enters as tensors, tensor operations transform that data, and the resulting computation contributes to a prediction. A loss function then measures the result.
One Training Cycle
Trace the state of a neural network during one conceptual training cycle.
Current state: The network begins with its current parameters. These parameters are used by the tensor computations.
Forward computation: Organized input data enters as tensors. Tensor operations transform the data and produce a prediction.
Evaluation: A loss function measures the result of the prediction.
Backward information: Backpropagation computes gradients by chaining derivatives through the operations.
Update: A gradient-based optimization method, such as stochastic gradient descent, uses the gradients to update the model parameters.
Next state: The updated parameters affect later tensor computations, so the next cycle begins with a changed model state.
Training is a repeated state change: compute, evaluate, calculate gradients, update parameters, and use the updated parameters in later computations.
Backpropagation and Parameter Updates
After the loss has been evaluated, the network needs information for changing its parameters. Backpropagation is the algorithm used to compute gradients by chaining derivatives through the operations. The gradients are then used by gradient-based optimization to update the parameters while minimizing the loss. Stochastic gradient descent is identified as one gradient-based optimization method.
Saying that backpropagation directly changes the parameters.
The source identifies backpropagation as the algorithm that computes gradients. An optimization method uses those gradients to update parameters.
Fix:
Say that backpropagation computes gradient information and stochastic gradient descent or another optimization method uses that information for the update.Treating gradients as the loss itself.
The loss measures the result, while gradients are computed from the training process to provide information for changing parameters.
Fix:
Trace the order as prediction, loss evaluation, gradient computation, and parameter update.Describing training as a one-time computation.
Training is described as a repeated state change in which updated parameters affect later tensor computations.
Fix:
Include the repeated cycle of computation, evaluation, gradient calculation, parameter update, and later computation.
Practice the Complete Trace
Explain the following training trace in your own words: organized input data enters as tensors; tensor operations transform it; the network produces a prediction; a loss function evaluates the result; backpropagation computes gradients; stochastic gradient descent uses the gradients to update parameters; and the updated parameters affect later tensor computations.
Hints
- First separate data representation from parameter optimization.
- Identify which step produces the gradients.
- Identify which step changes the parameters.
- End by explaining why the next computation is different from the previous one.
What do you think happens?
A learner says, "The network learned as soon as tensor operations produced a prediction." Is that complete?
Reveal answer
Answer: No, because the loss, gradients, and parameter update are also part of training
Tensor operations transform the data and contribute to a prediction, but training also requires evaluating the result, computing gradients through backpropagation, and updating parameters with a gradient-based optimization method.
The Complete Mental Model
- Deep learning's recent popularity reflects the combined influence of hardware, data, algorithms, investment, and democratization.
- Scalars, vectors, matrices, and higher-dimensional structures are tensor representations with different dimensions and indexing arrangements.
- Element-wise operations and tensor dot calculate on values, broadcasting supports some operations involving different shapes, and reshaping changes organization.
- Backpropagation computes gradients by chaining derivatives through operations, while stochastic gradient descent is a gradient-based method for updating parameters.
- The complete training trace is organized input tensors, tensor operations, prediction, loss evaluation, gradient computation, parameter update, and later computation with updated parameters.
Key Takeaways
- Deep learning became more popular through the combined influence of hardware, data, algorithms, investment, and democratization.
- Tensors organize neural-network data from zero-dimensional scalars through vectors, matrices, and higher-dimensional structures.
- Tensor operations transform representations and values through element-wise operations, broadcasting, tensor dot, and reshaping.
- Training repeatedly connects tensor computation with loss evaluation, backpropagation, gradients, and parameter updates.
- Tensor operations describe how data is transformed, while gradient-based optimization describes how model parameters change.