Matrices and Arrays
A 3D tensor is a collection of matrices packed into a new array.
From One Matrix to Many
A matrix organizes numbers into rows and columns. The key idea behind higher-dimensional tensors is not that the numbers become unfamiliar. Instead, complete lower-dimensional structures are grouped together inside a new array. When several matrices are packed together, the result is a 3D tensor.
Packing Matrices into Layers
Imagine several matrices placed one behind another. Each matrix remains a complete collection of rows and columns, but the group now has an additional way to distinguish one matrix from another: its layer in the packed collection. This layered arrangement is a 3D tensor. It can also be visualized as a cube of numbers.
Reading a packed collection
Suppose three matrices are packed into one array. How can you locate one value in the resulting 3D tensor?
Choose a layer: First select which packed matrix you want to inspect.
Choose a row: Within that matrix, select a row.
Choose a column: Within the selected row, select a column to reach the value.
A 3D tensor can be read in two stages: select a matrix layer, then locate a value by its row and column.
The extra dimension does not erase the original matrix structure. It identifies which complete matrix is being accessed.
Grouping 3D Tensors
The same construction continues one level higher. If a 3D tensor is already a packed collection of matrices, then packing several 3D tensors into another array produces a 4D tensor. The new grouping keeps each 3D tensor intact while adding another way to distinguish one packed 3D structure from another.
Following the nesting
A collection contains several 3D tensors. What does one value represent inside the resulting 4D tensor?
Select a 3D tensor: Use the added grouping level to choose one of the packed 3D tensors.
Select a matrix layer: Inside that 3D tensor, choose one of its packed matrices.
Select a row and column: Inside the selected matrix, locate the value using its row and column.
A 4D tensor preserves the nested structure: choose a 3D tensor, then a matrix layer, then a row and column.
The Dimension Ladder
The source describes deep learning as commonly using tensors from 0D through 4D. A useful way to study that range is as a ladder of grouped structures: a single value, a one-dimensional collection, a matrix, a collection of matrices, and a collection of 3D tensors. Each step adds a way to organize complete structures from the previous step.
The important pattern is structural continuity. Moving upward does not require abandoning the lower-dimensional view. A 3D tensor is still understandable through its matrix layers, and a 4D tensor is still understandable through the 3D tensors it contains.
Why Deep Learning Uses Several Dimensions
Deep learning commonly uses tensors from 0D to 4D because different data arrangements require different levels of grouping. The useful question is not simply which dimension is largest. It is which lower-dimensional structures need to be grouped together to represent the data being processed.
When deciding how to describe data, identify the complete structures being grouped. If the units are matrices, the packed result is 3D. If the units are already 3D tensors, the packed result is 4D.
Video's Fifth Dimension
Video processing may require 5D tensors. This extends the same nesting principle beyond the 0D-to-4D range commonly used in deep learning. The source identifies video processing as a case where another level of organization may be needed; the exact dimensionality depends on how the video data is arranged.
Common Packing Mistakes
Treating a 3D tensor as one oversized matrix
A 3D tensor is a collection of matrices, so a layer-selection step is needed before using a matrix's row and column structure.
Fix:
Read it in two stages: choose a matrix layer, then locate the value by row and column.Assuming that a 4D tensor loses its 3D structure
A 4D tensor is formed by packing complete 3D tensors into an array.
Fix:
Read from the outside inward: select a 3D tensor, then one of its matrix layers, then a row and column.Choosing a dimension without asking what is being grouped
The useful dimension depends on how the data is organized and which complete structures need to be grouped.
Fix:
Name the structures being packed before naming the resulting tensor dimension.
Check Your Understanding
Explain, in your own words, why packing four matrices creates a 3D tensor rather than a larger 2D matrix. Then explain what additional structure appears when several of those 3D tensors are packed into a 4D tensor.
Hints
- Focus on the complete structures being grouped.
- For the 3D case, identify the matrix layer before the row and column.
- For the 4D case, identify the 3D tensor before the matrix layer.
What do you think happens?
Before reading the answer, predict the sequence needed to locate one value in a 4D tensor built by packing 3D tensors.
Reveal answer
Answer: Choose the packed 3D tensor, choose a matrix layer inside it, and then use the row and column within that matrix.
Each higher level preserves the complete lower-dimensional structure inside it.
Key Takeaways
- A 3D tensor is a collection of matrices packed into a new array.
- A 3D tensor can be understood by choosing a matrix layer and then using that matrix's row and column structure.
- A 4D tensor is formed by packing 3D tensors into another array.
- Deep learning commonly uses tensors from 0D through 4D because data may require progressively higher levels of grouping.
- Video processing may require 5D tensors, showing that the needed dimensionality depends on how the data is organized.
Key Takeaways
- Packing matrices along a new axis produces a 3D tensor.
- Packing 3D tensors along another axis produces a 4D tensor.
- Higher-dimensional tensors preserve the lower-dimensional structures inside them.
- Deep learning commonly uses tensors from 0D through 4D, while video processing may require 5D tensors.
- To determine the useful dimensionality, ask which complete lower-dimensional structures must be grouped together.