Introduction to Keras
Tensors provide the data representations used by neural networks, from 0D scalars through higher-dimensional forms.
The Data Journey
A neural network needs a way to hold the numerical data it works with. That container is a tensor. Tensors provide the data representations used by neural networks, beginning with a single value and extending to structures with many dimensions. As data moves through a neural network, its tensor representation can be transformed, processed by layers, and used during training.
A useful way to understand the topic is to trace four questions. How is the data represented? How can its representation be transformed? How does the network improve its parameters during training? How do layers and models organize these transformations? Keras is introduced after these ideas as a deep-learning framework for developing with models built from layers and tensor-based data.
Tensor Dimensions
A tensor is a container for data, almost always numerical data. Tensors generalize matrices to an arbitrary number of dimensions. In this terminology, a matrix is a two-dimensional tensor, while a tensor dimension is often called an axis.
The dimension count describes the structural form in which the data is organized. A scalar is a 0D tensor because it has no dimension. A vector is a 1D tensor. A matrix is a 2D tensor. Data can also be represented with 3D tensors and tensors with still higher dimensionality.
| Representation | Tensor description | What the dimension count tells you |
|---|---|---|
| Scalar | 0D tensor | A single numerical value |
| Vector | 1D tensor | A one-dimensional organization of values |
| Matrix | 2D tensor | A two-dimensional organization of values |
| Higher-dimensional tensor | 3D or more | An organization with three or more dimensions |
The source describes these forms as tensor representations used by neural networks.
The broad idea applies to several data types discussed in this topic. Vector data, timeseries or sequence data, image data, and video data can all be discussed as data tensors. Their dimensions do more than hold numbers: they organize the data for the neural network.
Transforming Tensor Data
Tensor operations manipulate either tensor values or tensor structure. The operations highlighted in this topic are element-wise operations, broadcasting, tensor dot operations, and reshaping. They are different ways of working with the numerical contents of tensors or with the way those contents are organized.
| Operation | Purpose in the tensor workflow |
|---|---|
| Element-wise operation | Works with corresponding tensor values |
| Broadcasting | Supports combining tensor data when the operation applies across a broader arrangement of values |
| Tensor dot | Provides a tensor operation for combining tensor data |
| Reshaping | Changes the organization or structure of tensor data |
These descriptions identify the distinct roles named in the source without treating them as interchangeable operations.
Following a Tensor Through Operations
Suppose a neural-network workflow starts with numerical data stored in a tensor. How can the data be changed without losing the idea that it remains tensor data?
Start with organized data: The data is represented as a tensor, so the numerical values have an organized dimensional structure.
Apply an element-wise operation: An element-wise operation works with corresponding values in the tensor data.
Use broadcasting when appropriate: Broadcasting is another way of applying tensor operations across an arrangement of values.
Use a tensor dot operation: A tensor dot operation combines tensor data through a named tensor operation.
Reshape the tensor: Reshaping changes the tensor's organization or structure while keeping the focus on tensor data.
The important distinction is between operations that work with values and reshaping, which works with tensor structure. All four operations belong to the broader process of transforming neural-network data.
Training by Improvement
Gradient-based optimization is the training engine of a neural network. The training sequence connects predictions, loss, derivatives, gradients, and stochastic gradient descent. The network first produces a prediction. Training evaluates the result through a loss. Derivative information is represented by a gradient, and stochastic gradient descent uses that information in optimization.
A gradient provides information about change in the context of the tensor operations used during training. Stochastic gradient descent then uses that information to update the model during optimization. The process is cyclical rather than one-time: an updated model can take part in another training step.
One Abstract Optimization Step
Trace one pass through the training process without calculating numerical values.
Represent: The neural network receives numerical data represented as tensors.
Predict: The network transforms the data and produces a prediction.
Evaluate: Training evaluates the prediction through a loss.
Differentiate: Derivative information is represented by a gradient.
Optimize: Stochastic gradient descent uses the gradient during optimization and updates the model.
Repeat: The updated model can participate in another training step, making the process cyclical.
Gradient-based optimization connects the network's prediction to a repeated process of evaluation and parameter improvement.
Layers, Models, and Keras
Layers are the building blocks of deep-learning models. A model is a network of layers. Keras is a deep-learning framework whose practical role is to support development with deep-learning models.
This gives neural-network architecture two useful levels of description. An individual layer is a component that data can move through. A model organizes layers into a network. The sequence of layers constitutes the model's network, so the model is more than a single layer.
Keras fits into the overall workflow after the core ideas have been identified: tensors represent data, tensor operations transform it, gradient-based optimization supports training, and layers organize transformations into models. Keras is the framework used to develop with this deep-learning model structure.
Treating a layer and a model as the same thing.
The source distinguishes layers as building blocks and models as networks of layers.
Fix:
Use layer for an individual building block and model for the organized network of layers.Thinking Keras replaces the tensor representation.
Tensors remain the data representations used by neural networks; Keras is introduced as a deep-learning framework.
Fix:
Keep the roles separate: tensors represent data, while Keras supports development with deep-learning models.Describing training as a single prediction step.
The training process continues through loss evaluation, gradient information, stochastic gradient descent, and another training step.
Fix:
Trace the complete cycle from prediction to loss, gradient, optimization, and repetition.
Practice and Summary
Explain the following workflow in your own words: numerical data is represented as a tensor, transformed through operations and layers, used to produce a prediction, evaluated through a loss, and involved in a gradient-based optimization cycle. In your explanation, use the terms scalar, vector, matrix, axis, model, and Keras at least once.
Hints
- Start by distinguishing the broad category tensor from the specific form matrix.
- Remember that a dimension is often called an axis.
- Separate the roles of prediction, loss, gradient, and stochastic gradient descent.
- End by explaining how layers form a model and where Keras fits.
Calling every tensor a matrix.
Fix:
A matrix is a 2D tensor; tensors may have zero, one, two, three, or more dimensions.Using axis as if it meant a separate object from a dimension.
Fix:
In tensor terminology, a dimension is often called an axis.Listing tensor operations without distinguishing their purposes.
Fix:
Element-wise operations, broadcasting, tensor dot, and reshaping are different ways to manipulate tensor values or tensor structure.Saying that stochastic gradient descent creates the prediction.
Fix:
The network produces the prediction; stochastic gradient descent uses gradient information during optimization.
- A tensor is a container for numerical data and the basic data structure used by current machine-learning systems.
- Scalars, vectors, matrices, and higher-dimensional tensors are tensor representations with different numbers of dimensions, or axes.
- Element-wise operations, broadcasting, tensor dot operations, and reshaping manipulate tensor values or structure in different ways.
- Gradient-based optimization connects prediction, loss, derivatives, gradients, and stochastic gradient descent in a repeating training cycle.
- Layers are model building blocks, models are networks of layers, and Keras is a deep-learning framework for developing with these models.
Key Takeaways
- Tensors organize the numerical data used by neural networks, from 0D scalars through higher-dimensional forms.
- A matrix is a 2D tensor, and a tensor dimension is often called an axis.
- Tensor operations transform values or structure through element-wise operations, broadcasting, tensor dot operations, and reshaping.
- Training follows a cycle in which predictions are evaluated through loss, gradients guide stochastic gradient descent, and the model is updated.
- Layers form models, while Keras provides a framework for developing deep-learning models.