Concepts / Matrices and Arrays

Matrices and Arrays

A 3D tensor is a collection of matrices packed into a new array.

  • Programming

From One Matrix to Many

A matrix organizes numbers into rows and columns. The key idea behind higher-dimensional tensors is not that the numbers become unfamiliar. Instead, complete lower-dimensional structures are grouped together inside a new array. When several matrices are packed together, the result is a 3D tensor.

Packing Matrices into Layers

Imagine several matrices placed one behind another. Each matrix remains a complete collection of rows and columns, but the group now has an additional way to distinguish one matrix from another: its layer in the packed collection. This layered arrangement is a 3D tensor. It can also be visualized as a cube of numbers.

packed as a layerpacked as a layerpacked as a layerMatrix 1rows and columns3D tensorpacked layersMatrix 2rows and columnsMatrix 3rows and columns
How do multiple 2D matrices stack along a new axis to form a 3D tensor?

Reading a packed collection

Suppose three matrices are packed into one array. How can you locate one value in the resulting 3D tensor?

Choose a layer: First select which packed matrix you want to inspect.

Choose a row: Within that matrix, select a row.

Choose a column: Within the selected row, select a column to reach the value.

A 3D tensor can be read in two stages: select a matrix layer, then locate a value by its row and column.

The extra dimension does not erase the original matrix structure. It identifies which complete matrix is being accessed.

Grouping 3D Tensors

The same construction continues one level higher. If a 3D tensor is already a packed collection of matrices, then packing several 3D tensors into another array produces a 4D tensor. The new grouping keeps each 3D tensor intact while adding another way to distinguish one packed 3D structure from another.

grouped insidegrouped insidegrouped inside3D tensor 1packed matrices4D tensorpacked 3D tensors3D tensor 2packed matrices3D tensor 3packed matrices
What happens when multiple 3D tensors are grouped along another axis to produce a 4D tensor?

Following the nesting

A collection contains several 3D tensors. What does one value represent inside the resulting 4D tensor?

Select a 3D tensor: Use the added grouping level to choose one of the packed 3D tensors.

Select a matrix layer: Inside that 3D tensor, choose one of its packed matrices.

Select a row and column: Inside the selected matrix, locate the value using its row and column.

A 4D tensor preserves the nested structure: choose a 3D tensor, then a matrix layer, then a row and column.

The Dimension Ladder

The source describes deep learning as commonly using tensors from 0D through 4D. A useful way to study that range is as a ladder of grouped structures: a single value, a one-dimensional collection, a matrix, a collection of matrices, and a collection of 3D tensors. Each step adds a way to organize complete structures from the previous step.

adds an organizing axisadds an organizing axispacks matricespacks 3D tensorsScalar0DVector1DMatrix2D3D tensorpacked matrices4D tensorpacked 3D tensors
How do scalar, vector, matrix, 3D tensor, and 4D tensor structures differ as dimensions are added?

The important pattern is structural continuity. Moving upward does not require abandoning the lower-dimensional view. A 3D tensor is still understandable through its matrix layers, and a 4D tensor is still understandable through the 3D tensors it contains.

Why Deep Learning Uses Several Dimensions

Deep learning commonly uses tensors from 0D to 4D because different data arrangements require different levels of grouping. The useful question is not simply which dimension is largest. It is which lower-dimensional structures need to be grouped together to represent the data being processed.

may requiremay requiremay requiremay requiremay requireOrganized datadifferent structures0D tensorsingle value1D tensorone collection2D tensorrows and columns3D tensorpacked matrices4D tensorpacked 3D tensors
How do different tensor dimensions correspond to the kinds of data commonly processed in deep learning?

When deciding how to describe data, identify the complete structures being grouped. If the units are matrices, the packed result is 3D. If the units are already 3D tensors, the packed result is 4D.

Video's Fifth Dimension

Video processing may require 5D tensors. This extends the same nesting principle beyond the 0D-to-4D range commonly used in deep learning. The source identifies video processing as a case where another level of organization may be needed; the exact dimensionality depends on how the video data is arranged.

organized along one axisorganized along one axisorganized along one axisorganized along one axisorganized along one axisBatchesgrouped videos5D tensorvideo arrangementFramesvideo sequenceHeightframe rowsWidthframe columnsChannelsframe information
How are frames, height, width, channels, and batches represented together in a 5D tensor for video processing?

Common Packing Mistakes

  • Treating a 3D tensor as one oversized matrix

    A 3D tensor is a collection of matrices, so a layer-selection step is needed before using a matrix's row and column structure.

    Fix: Read it in two stages: choose a matrix layer, then locate the value by row and column.

  • Assuming that a 4D tensor loses its 3D structure

    A 4D tensor is formed by packing complete 3D tensors into an array.

    Fix: Read from the outside inward: select a 3D tensor, then one of its matrix layers, then a row and column.

  • Choosing a dimension without asking what is being grouped

    The useful dimension depends on how the data is organized and which complete structures need to be grouped.

    Fix: Name the structures being packed before naming the resulting tensor dimension.

Check Your Understanding

MEDIUM

Explain, in your own words, why packing four matrices creates a 3D tensor rather than a larger 2D matrix. Then explain what additional structure appears when several of those 3D tensors are packed into a 4D tensor.

Hints
  • Focus on the complete structures being grouped.
  • For the 3D case, identify the matrix layer before the row and column.
  • For the 4D case, identify the 3D tensor before the matrix layer.

What do you think happens?

Before reading the answer, predict the sequence needed to locate one value in a 4D tensor built by packing 3D tensors.

Reveal answer

Answer: Choose the packed 3D tensor, choose a matrix layer inside it, and then use the row and column within that matrix.

Each higher level preserves the complete lower-dimensional structure inside it.

Key Takeaways

  1. A 3D tensor is a collection of matrices packed into a new array.
  2. A 3D tensor can be understood by choosing a matrix layer and then using that matrix's row and column structure.
  3. A 4D tensor is formed by packing 3D tensors into another array.
  4. Deep learning commonly uses tensors from 0D through 4D because data may require progressively higher levels of grouping.
  5. Video processing may require 5D tensors, showing that the needed dimensionality depends on how the data is organized.

Key Takeaways

  • Packing matrices along a new axis produces a 3D tensor.
  • Packing 3D tensors along another axis produces a 4D tensor.
  • Higher-dimensional tensors preserve the lower-dimensional structures inside them.
  • Deep learning commonly uses tensors from 0D through 4D, while video processing may require 5D tensors.
  • To determine the useful dimensionality, ask which complete lower-dimensional structures must be grouped together.