Concepts / Introduction to Keras

Introduction to Keras

Tensors provide the data representations used by neural networks, from 0D scalars through higher-dimensional forms.

  • Programming

The Data Journey

A neural network needs a way to hold the numerical data it works with. That container is a tensor. Tensors provide the data representations used by neural networks, beginning with a single value and extending to structures with many dimensions. As data moves through a neural network, its tensor representation can be transformed, processed by layers, and used during training.

A useful way to understand the topic is to trace four questions. How is the data represented? How can its representation be transformed? How does the network improve its parameters during training? How do layers and models organize these transformations? Keras is introduced after these ideas as a deep-learning framework for developing with models built from layers and tensor-based data.

more dimensionsmore dimensionsmore dimensionsScalar0D tensorVector1D tensorMatrix2D tensorHigher-dimensionaltensor3D or more
What changes as numerical data moves from a 0D scalar to 1D, 2D, and higher-dimensional tensor representations?

Tensor Dimensions

A tensor is a container for data, almost always numerical data. Tensors generalize matrices to an arbitrary number of dimensions. In this terminology, a matrix is a two-dimensional tensor, while a tensor dimension is often called an axis.

The dimension count describes the structural form in which the data is organized. A scalar is a 0D tensor because it has no dimension. A vector is a 1D tensor. A matrix is a 2D tensor. Data can also be represented with 3D tensors and tensors with still higher dimensionality.

RepresentationTensor descriptionWhat the dimension count tells you
Scalar0D tensorA single numerical value
Vector1D tensorA one-dimensional organization of values
Matrix2D tensorA two-dimensional organization of values
Higher-dimensional tensor3D or moreAn organization with three or more dimensions

The source describes these forms as tensor representations used by neural networks.

The broad idea applies to several data types discussed in this topic. Vector data, timeseries or sequence data, image data, and video data can all be discussed as data tensors. Their dimensions do more than hold numbers: they organize the data for the neural network.

hashasdescribed byTensornumerical dataAxis 0one dimensionAxis 1another dimensionShapedimension organization
How do tensor axes, dimensions, and shape describe the position and organization of each value?

Transforming Tensor Data

Tensor operations manipulate either tensor values or tensor structure. The operations highlighted in this topic are element-wise operations, broadcasting, tensor dot operations, and reshaping. They are different ways of working with the numerical contents of tensors or with the way those contents are organized.

OperationPurpose in the tensor workflow
Element-wise operationWorks with corresponding tensor values
BroadcastingSupports combining tensor data when the operation applies across a broader arrangement of values
Tensor dotProvides a tensor operation for combining tensor data
ReshapingChanges the organization or structure of tensor data

These descriptions identify the distinct roles named in the source without treating them as interchangeable operations.

Following a Tensor Through Operations

Suppose a neural-network workflow starts with numerical data stored in a tensor. How can the data be changed without losing the idea that it remains tensor data?

Start with organized data: The data is represented as a tensor, so the numerical values have an organized dimensional structure.

Apply an element-wise operation: An element-wise operation works with corresponding values in the tensor data.

Use broadcasting when appropriate: Broadcasting is another way of applying tensor operations across an arrangement of values.

Use a tensor dot operation: A tensor dot operation combines tensor data through a named tensor operation.

Reshape the tensor: Reshaping changes the tensor's organization or structure while keeping the focus on tensor data.

The important distinction is between operations that work with values and reshaping, which works with tensor structure. All four operations belong to the broader process of transforming neural-network data.

applyapplyapplyreorganizeTensor datavalues and structureElement-wisecorresponding valuesBroadcastingbroader arrangementTensor dottensor combinationReshaped tensorchanged organization
How do tensor values and shapes change when tensors are combined element by element, broadcast, multiplied with a dot operation, or reshaped?

Training by Improvement

Gradient-based optimization is the training engine of a neural network. The training sequence connects predictions, loss, derivatives, gradients, and stochastic gradient descent. The network first produces a prediction. Training evaluates the result through a loss. Derivative information is represented by a gradient, and stochastic gradient descent uses that information in optimization.

A gradient provides information about change in the context of the tensor operations used during training. Stochastic gradient descent then uses that information to update the model during optimization. The process is cyclical rather than one-time: an updated model can take part in another training step.

network processesevaluated byderivative informationguidesupdatesrepeat trainingTraining datatensor dataPredictionnetwork outputLosstraining evaluationGradientderivative informationStochastic gradientdescentoptimization updateUpdated modelnext training step
How do predictions produce a loss, how does the gradient determine parameter changes, and how does training repeat this process across batches?

One Abstract Optimization Step

Trace one pass through the training process without calculating numerical values.

Represent: The neural network receives numerical data represented as tensors.

Predict: The network transforms the data and produces a prediction.

Evaluate: Training evaluates the prediction through a loss.

Differentiate: Derivative information is represented by a gradient.

Optimize: Stochastic gradient descent uses the gradient during optimization and updates the model.

Repeat: The updated model can participate in another training step, making the process cyclical.

Gradient-based optimization connects the network's prediction to a repeated process of evaluation and parameter improvement.

Layers, Models, and Keras

Layers are the building blocks of deep-learning models. A model is a network of layers. Keras is a deep-learning framework whose practical role is to support development with deep-learning models.

This gives neural-network architecture two useful levels of description. An individual layer is a component that data can move through. A model organizes layers into a network. The sequence of layers constitutes the model's network, so the model is more than a single layer.

data moves throughdata moves throughorganized asTensor datamodel inputLayer 1model componentLayer 2model componentModelnetwork of layers
How does data flow from one Keras layer to the next, and how do connected layers form a complete model?

Keras fits into the overall workflow after the core ideas have been identified: tensors represent data, tensor operations transform it, gradient-based optimization supports training, and layers organize transformations into models. Keras is the framework used to develop with this deep-learning model structure.

  • Treating a layer and a model as the same thing.

    The source distinguishes layers as building blocks and models as networks of layers.

    Fix: Use layer for an individual building block and model for the organized network of layers.

  • Thinking Keras replaces the tensor representation.

    Tensors remain the data representations used by neural networks; Keras is introduced as a deep-learning framework.

    Fix: Keep the roles separate: tensors represent data, while Keras supports development with deep-learning models.

  • Describing training as a single prediction step.

    The training process continues through loss evaluation, gradient information, stochastic gradient descent, and another training step.

    Fix: Trace the complete cycle from prediction to loss, gradient, optimization, and repetition.

Practice and Summary

MEDIUM

Explain the following workflow in your own words: numerical data is represented as a tensor, transformed through operations and layers, used to produce a prediction, evaluated through a loss, and involved in a gradient-based optimization cycle. In your explanation, use the terms scalar, vector, matrix, axis, model, and Keras at least once.

Hints
  • Start by distinguishing the broad category tensor from the specific form matrix.
  • Remember that a dimension is often called an axis.
  • Separate the roles of prediction, loss, gradient, and stochastic gradient descent.
  • End by explaining how layers form a model and where Keras fits.
  • Calling every tensor a matrix.

    Fix: A matrix is a 2D tensor; tensors may have zero, one, two, three, or more dimensions.

  • Using axis as if it meant a separate object from a dimension.

    Fix: In tensor terminology, a dimension is often called an axis.

  • Listing tensor operations without distinguishing their purposes.

    Fix: Element-wise operations, broadcasting, tensor dot, and reshaping are different ways to manipulate tensor values or tensor structure.

  • Saying that stochastic gradient descent creates the prediction.

    Fix: The network produces the prediction; stochastic gradient descent uses gradient information during optimization.

  1. A tensor is a container for numerical data and the basic data structure used by current machine-learning systems.
  2. Scalars, vectors, matrices, and higher-dimensional tensors are tensor representations with different numbers of dimensions, or axes.
  3. Element-wise operations, broadcasting, tensor dot operations, and reshaping manipulate tensor values or structure in different ways.
  4. Gradient-based optimization connects prediction, loss, derivatives, gradients, and stochastic gradient descent in a repeating training cycle.
  5. Layers are model building blocks, models are networks of layers, and Keras is a deep-learning framework for developing with these models.

Key Takeaways

  • Tensors organize the numerical data used by neural networks, from 0D scalars through higher-dimensional forms.
  • A matrix is a 2D tensor, and a tensor dimension is often called an axis.
  • Tensor operations transform values or structure through element-wise operations, broadcasting, tensor dot operations, and reshaping.
  • Training follows a cycle in which predictions are evaluated through loss, gradients guide stochastic gradient descent, and the model is updated.
  • Layers form models, while Keras provides a framework for developing deep-learning models.