Concepts / Getting started with neural networks

Getting started with neural networks

Deep learning became more popular through the combined influence of hardware, data, algorithms, investment, and democratization.

  • Programming

Why deep learning became practical

Deep learning became especially prominent through several factors working together rather than through one isolated breakthrough. Advances in hardware made larger computations more practical. Greater data availability supplied more material for learning. Improvements in algorithms made learning methods more effective. Investment increased activity in the field, while democratization made deep learning more accessible to a wider group of practitioners. These influences form the context for understanding neural networks today.

contributedcontributedcontributedcontributedcontributedHardwareDeep learningprominenceDataAlgorithmsInvestmentDemocratization
How did several contributing factors combine to make deep learning practical and widely accessible?

The useful mental model is a combination: the surrounding ecosystem helped deep learning become prominent, while tensors and optimization explain how a neural network computes and learns.

From values to tensors

A tensor organizes values into dimensions and positions. A scalar is a zero-dimensional tensor containing a single value. A vector is one-dimensional and can be understood as an ordered collection of values. A matrix is two-dimensional, organizing values across two dimensions. Higher-dimensional tensors extend this arrangement to additional dimensions. The dimensions and indices describe where values are arranged, and data batches are part of neural-network data representation.

Reading tensor representations

Classify four illustrative data arrangements by their number of dimensions.

Scalar: A single value has no organizing dimension, so it is a zero-dimensional tensor.

Vector: An ordered row of values uses one dimension and positions values along that dimension.

Matrix: Values arranged by rows and columns use two dimensions.

Higher-dimensional tensor: Adding another organizing dimension produces a tensor with more than two dimensions.

The number of dimensions describes the arrangement of values, while indices identify positions within that arrangement.

adds an organizing dimensionadds an organizing dimensionadds dimensions70 dimensions[2, 4, 6]1 dimension[[1, 2], [3, 4]]2 dimensionstensormore than 2 dimensions
What does each tensor shape look like, and how do scalars, vectors, matrices, and higher-dimensional tensors differ in their dimensions and positions?

Transforming tensor data

Tensor operations are the computational steps that manipulate both data representations and values inside a neural network. Four important categories are element-wise operations, broadcasting, tensor dot, and reshaping. Element-wise operations calculate on corresponding values. Broadcasting addresses operations involving tensors with different shapes when the operation's rules allow those tensors to work together. Tensor dot is a dot operation over tensor values. Reshaping changes how values are organized in the tensor.

OperationWhat it changesWhat to inspect
Element-wise operationCalculations on corresponding tensor valuesWhich values are paired
BroadcastingPermits an operation between tensors with different shapes when the rules allow itWhether the shapes are compatible
Tensor dotPerforms a dot-style calculation over tensor valuesHow the operation combines the tensors
ReshapingChanges the organization or shape of the tensorHow positions are rearranged

The four tensor-operation categories identified in the source

Separating value changes from shape changes

Consider an illustrative tensor arranged as two rows of two values. Identify what kind of change each operation represents.

Element-wise operation: A calculation is applied to corresponding values, so the values can change while the arrangement can remain the same.

Broadcasting: A differently shaped tensor participates in an operation when the operation's shape rules allow it. The important question is compatibility, not simply whether the shapes look identical.

Tensor dot: The tensor values are combined through a dot-style calculation. This is a value-producing operation whose result depends on the tensors and the operation being used.

Reshaping: The organization changes. The same tensor data is represented with a different arrangement of dimensions.

Element-wise operations and tensor dot focus on calculations over values, broadcasting enables certain different-shape calculations, and reshaping changes organization.

operationoperationoperationoperationTensorvalues and shapeElement-wise resultcorresponding valuescalculatedBroadcasted resultdifferent shapes permittedwhen compatibleTensor dot resultdot-style value calculationReshaped tensororganization changes
What changes in tensor values and shape when common tensor operations are applied?

The forward-to-update cycle

Neural-network learning is a repeated state change. Organized input data enters as tensors. The network uses its current parameters in tensor operations to produce a result. A loss function evaluates that result. Backpropagation then computes gradients by chaining derivatives through the operations. Finally, a gradient-based optimization method updates the parameters. Those updated parameters affect later tensor computations.

entersproducesevaluated bydrives computation ofguidesaffects later computationsInput datatensorTensor operationscurrent parametersPredictionLossGradientsbackpropagationUpdated parametersoptimization
How does input data move through tensor operations to produce a prediction, loss, gradients, and updated model parameters?

Tensors and optimization answer different questions. Tensor operations describe how data is represented and transformed. Gradient-based optimization describes how the parameters used by those transformations change.

Backpropagation and parameter updates

A gradient provides information used to change model parameters while minimizing a loss function. Backpropagation is the algorithm used to compute gradients by chaining derivatives through the operations. This means the error-related signal is traced through the connected computations so that gradients can be associated with the parameters involved.

Stochastic gradient descent is a gradient-based optimization method. In the training cycle, it uses the computed gradients to update parameters, after which the network performs later tensor computations with the changed parameters. The cycle is repeated rather than completed by a single operation.

backpropagates throughcontinues throughcontributes toguides changes toLossOperation Bconnected computationOperation Aconnected computationGradientschained derivativesParameters
How does the error signal move backward through connected operations to determine how parameters should change?

One training step at a time

The training process can be understood as an iteration. A batch of organized data is processed with the current parameters. Tensor operations produce the network's result, and a loss function evaluates it. Backpropagation computes gradients. Stochastic gradient descent updates the parameters. The next iteration therefore begins with a changed model state.

processed bycontributes toused to computeguidesleads torepeats withData batchtensor representationTensor computationcurrent parametersLossGradientsbackpropagationParameter updatestochastic gradient descentNext iterationchanged parameters
What happens during each training step as a batch is processed, gradients are computed, and parameters are updated repeatedly?

What do you think happens?

After gradients have been computed, what happens before the next tensor computation?

  • The parameters are updated by a gradient-based optimization method
  • The tensor representation is automatically discarded
  • Backpropagation replaces the loss function
  • No state changes occur
Reveal answer

Answer: The parameters are updated by a gradient-based optimization method.

The training cycle uses gradients to update parameters. Those updated parameters then affect later tensor computations.

Mistakes to avoid

  • Treating a tensor as only a list of values

    Tensor dimensions and indices describe how values are arranged.

    Fix: Track both the values and the dimensions that organize them.

  • Confusing reshaping with a value calculation

    Reshaping changes the organization of the tensor, while element-wise operations and tensor dot perform calculations on values.

    Fix: Ask whether the operation changes organization, values, or both.

  • Assuming broadcasting always works

    Broadcasting addresses different-shape operations only when the relevant rules allow them to work together.

    Fix: Check shape compatibility for the specific operation.

  • Treating backpropagation and optimization as identical

    Backpropagation computes gradients, while an optimization method uses those gradients to update parameters.

    Fix: Describe the sequence as compute gradients, then update parameters.

  • Thinking one operation is the whole learning process

    Training is a repeated process involving tensor computation, loss evaluation, gradient computation, and parameter updates.

    Fix: Trace the complete cycle and include the changed parameters in the next iteration.

Practice the complete trace

MEDIUM

Describe the path from a batch of input data to the next training iteration. Your answer should name the tensor representation, tensor operations, prediction, loss, backpropagation, gradients, stochastic gradient descent, and updated parameters.

Hints
  • Begin with organized input data represented as tensors.
  • Separate the forward computation and loss evaluation from the backward gradient computation.
  • End by explaining why the updated parameters affect later tensor operations.

A complete verbal trace

Explain one abstract training iteration without using a particular neural-network architecture or numerical values.

Represent: Organize the input batch as tensor data with dimensions and positions.

Transform: Apply tensor operations using the network's current parameters to produce a result.

Evaluate: Use a loss function to measure the result.

Differentiate: Use backpropagation to chain derivatives through the operations and compute gradients.

Update: Use stochastic gradient descent, a gradient-based optimization method, to update the parameters.

Repeat: Use the changed parameters in later tensor computations.

Learning connects tensor representation and transformation with gradient-based parameter updates in a repeated cycle.

Key takeaways

  1. Deep learning's recent prominence reflects the combined influence of hardware, data, algorithms, investment, and democratization.
  2. Scalars, vectors, matrices, and higher-dimensional tensors organize values across zero, one, two, or more dimensions.
  3. Element-wise operations, broadcasting, tensor dot, and reshaping manipulate tensor values or organization in different ways.
  4. Backpropagation computes gradients by chaining derivatives through operations, while stochastic gradient descent uses gradients to update parameters.
  5. The central training trace is tensor input, tensor operations, loss evaluation, gradient computation, parameter update, and later computation with the updated parameters.

Key Takeaways

  • Deep learning became prominent through the combined effects of hardware, data, algorithms, investment, and democratization.
  • Tensors organize neural-network data from scalars through higher-dimensional structures.
  • Tensor operations transform values or representations, while gradient-based optimization changes model parameters.
  • Backpropagation computes gradients and stochastic gradient descent applies them through repeated parameter updates.
  • Training is a connected cycle: tensors enter, operations produce a result, loss is evaluated, gradients are computed, and parameters are updated.