Data Representations for Neural Networks
Tensors provide the multi-dimensional representations used by neural networks.
From Raw Data to Learnable Structure
A neural network cannot learn directly from raw sound, images, text, or an unprepared table. Preprocessing changes raw data into numerical forms that the network can accept and learn from. The central representation is the tensor: a multi-dimensional array used to organize neural-network data.
Tensor Ranks and Axes
A tensor may be a scalar with no dimensions, a vector with one dimension, a matrix with two dimensions, or a higher-dimensional tensor with more dimensions. Dimensions and axes organize the values so that neural-network operations can act on them. The representation is not merely a container: its organization determines how values can be combined and transformed.
| Representation | Number of dimensions | What it organizes |
|---|---|---|
| Scalar | 0 | A single value |
| Vector | 1 | A one-dimensional sequence of values |
| Matrix | 2 | Values organized across two dimensions |
| Higher-dimensional tensor | More than 2 | Values organized across multiple dimensions |
The source distinguishes tensor representations by the number of dimensions they contain.
Classifying Tensor Representations
Classify three hypothetical neural-network inputs: one value, a list of values, and values arranged across two dimensions.
One value: A representation containing one value is a scalar because it has no dimensions.
A list of values: A one-dimensional sequence is a vector.
Two organized dimensions: Values arranged across two dimensions form a matrix.
The number of dimensions identifies whether the representation is a scalar, vector, matrix, or higher-dimensional tensor.
Operations That Change Tensors
Tensor operations are the transformations that let a neural network manipulate its representations and perform computations. Element-wise operations act on corresponding values. Broadcasting supports computation across dimensions when a value or smaller representation is applied across a larger one. A tensor dot operation combines values across dimensions. Reshaping changes the organization of values without being described as a new set of values. These roles should be kept distinct: some operations primarily change values, some combine values, and reshaping changes arrangement.
Why Learning Needs Gradients
Tensor operations provide the computation performed by a neural network, but computation alone does not explain learning. Learning uses gradient-based optimization to adjust the network's parameters and minimize its loss function. A gradient is described as derivative information from a tensor operation. It indicates how the computation changes in relation to the parameters, giving optimization a direction to use.
Stochastic gradient descent is one method used in this optimization process. Backpropagation chains derivatives so that the effects of tensor operations contribute to parameter updates. The connection is important: tensor operations produce the computation, their derivatives provide information about that computation, and stochastic gradient descent uses that information to move parameters toward a lower-loss region.
Preparing Values Before Training
Vectorization produces tensors from sources such as text, images, and tabular data. This step converts raw inputs and targets into numerical representations that neural-network computations can process. Both inputs and targets must be represented as tensors because the network's computations and learning procedures operate on tensor representations.
Neural-network data should usually contain small, relatively homogeneous values. Normalization keeps values small and helps avoid problems caused by incompatible feature ranges. Two related preparation ideas must be distinguished. Scaling places values into a small range. Standardization changes a feature so that its mean is 0 and its standard deviation is 1. Both address the relationship between feature magnitudes, but standardization specifically targets those two statistical properties.
A missing value may be represented as 0 only when 0 does not already have a meaningful interpretation for that feature. If missing entries are consistently represented this way during training, the network can learn that 0 signals missing data and can learn to ignore that value. If 0 is a meaningful value for the feature, using it as the missing marker would confuse two different situations.
Common Representation Mistakes
Treating raw data as ready for a neural network
The network requires numerical tensor representations.
Fix:
Preprocess the source through vectorization and other appropriate preparation steps.Calling every tensor operation a reshape
Element-wise operations, broadcasting, tensor dot operations, and reshaping have different roles.
Fix:
State whether the operation changes corresponding values, applies across dimensions, combines dimensions, or changes organization.Using a meaningful zero as the missing-value marker
The network cannot distinguish the two meanings from that representation.
Fix:
Use 0 as a missing marker only when 0 has no meaningful interpretation for the feature, and use the same convention consistently.Ignoring incompatible feature ranges
Normalization helps avoid problems caused by incompatible feature ranges.
Fix:
Prepare values so they are small and relatively homogeneous, using scaling or standardization as appropriate.
Check Your Understanding
A dataset contains text, image data, and a table with one feature that sometimes has no recorded value. Describe the preparation path before training. Identify the tensor representation involved, name the operation that changes organization without changing the set of values, distinguish scaling from standardization, and explain when 0 could safely represent the missing feature value.
Hints
- Start with vectorization: raw sources must become numerical tensors.
- A scalar, vector, matrix, or higher-dimensional tensor is identified by its number of dimensions.
- Reshaping changes organization, while other tensor operations may change or combine values.
- Scaling targets a small range; standardization targets a mean of 0 and a standard deviation of 1.
- Use 0 for missing data only when 0 is not already meaningful for that feature.
- A neural network receives and transforms tensors rather than raw data. Scalars, vectors, matrices, and higher-dimensional tensors organize values by their dimensions. Element-wise operations, broadcasting, tensor dot operations, and reshaping perform different kinds of transformations. Tensor computations produce the derivatives used by backpropagation and stochastic gradient descent to update parameters and minimize loss. Vectorization creates tensor inputs and targets, while normalization helps keep values small and relatively homogeneous. Scaling creates a small range, standardization targets a mean of 0 and a standard deviation of 1, and missing values require a marker that cannot be confused with a meaningful data value.
Key Takeaways
- Tensors represent neural-network data as scalars, vectors, matrices, or higher-dimensional arrays.
- Element-wise operations, broadcasting, tensor dot operations, and reshaping differ in how they manipulate values and organization.
- Gradients connect tensor computations to parameter updates performed by stochastic gradient descent.
- Vectorization converts raw sources into tensors, and normalization helps keep values small and relatively homogeneous.
- Missing values should use 0 as a marker only when 0 has no meaningful interpretation for that feature.