Feature extraction
Preprocessing changes raw data into numerical forms that neural networks can use.
From raw data to learnable input
A neural network cannot learn directly from raw sound, images, text, or an unprepared table. Before learning begins, preprocessing changes the data into numerical forms that the network can accept. This preparation includes vectorization, normalization, and decisions about missing values. Inputs and targets are represented as tensors because vectorization produces tensor values, and tensors provide the numerical structure used by the network.
Vectorization means changing a source such as text, an image, or tabular data into numerical values. Those values are organized into tensors before they are passed into the network. The same principle applies to targets: the expected answers must also be represented in a numerical tensor form that the learning process can use.
Why value ranges matter
Neural-network data should usually contain small, relatively homogeneous values. When different features use incompatible ranges, learning can encounter problems. Normalization addresses this by keeping values small and reducing the mismatch between feature ranges.
| Preparation | What it does | When to recognize it |
|---|---|---|
| Scaling | Maps values into a small range | Useful when the goal is to keep numerical values small |
| Standardizing | Transforms a feature to a mean of 0 and a standard deviation of 1 | Useful when describing each feature relative to its mean and spread |
Missing values without ambiguity
A dataset may lack a value for a feature in some training or test records. One safe strategy is to represent a missing value as 0, but only when 0 does not already have a meaningful interpretation for that feature. If 0 is a genuine value, using it as the missing marker makes two different situations look identical to the network.
Choosing a missing-value representation
A feature is absent in some records, but the value 0 is already meaningful for that feature.
Check the meaning of zero: Because 0 already represents a real value, it cannot safely serve as the missing marker.
Separate the meanings: Choose a representation that lets the network distinguish absence from a genuine zero.
Use it consistently: Represent missing entries consistently during training so the network can learn that the marker signals missing data.
Missing data is encoded separately from a meaningful value of zero.
Before choosing 0 as a missing marker, ask whether 0 already means something in the feature. If it does, use a representation that preserves the distinction between missing and present.
A pretrained network as a starting point
A pretrained network is a saved network trained previously on a large dataset, such as ImageNet. Its learned representations can be reused because the network has already learned visual patterns that may also help with a new image problem.
Feature extraction begins before training a classifier for the new image problem. You keep the part of the pretrained network that produces visual representations, pass the new images through that part, and train a separate classifier from scratch on the resulting representations.
The convolutional base and classifier
The data path has two different jobs. The convolutional base transforms an image into learned visual features. The new classifier interprets those features for the new set of classes. In the feature-extraction configuration, the original classification section is excluded and a new classifier is placed after the retained convolutional and pooling layers.
Depth and feature reusability
The depth of a convolutional layer affects how reusable its features are. Earlier convolutional layers tend to learn local visual patterns such as edges, colors, and textures. Deeper layers tend to represent more abstract patterns, such as a cat ear or a dog eye. Earlier features are generally more reusable, while deeper features may be more tied to the original task.
The best reuse boundary depends on how different the new dataset is from the original one. If the datasets differ substantially, the first few convolutional layers may be more appropriate because their features are more general. Higher layers encode increasingly abstract concepts and may be less reusable for a very different task.
Selecting a reuse boundary
A new image dataset differs substantially from the dataset used to train the pretrained network.
Compare the datasets: A substantial difference suggests that higher-level features may be tied too closely to the original image problem.
Prefer general representations: Earlier convolutional layers capture local patterns such as edges, colors, and textures.
Place the new classifier: Use the selected convolutional output as the input to a new classifier for the new classes.
Earlier convolutional features may provide a safer reuse boundary when the new dataset is very different.
Common reuse mistakes
Keeping the original classifier for a new set of classes.
The original classifier is tied to the classes from the first training task.
Fix:
Reuse the convolutional base and place a new classifier after its output.Treating the whole pretrained network as equally reusable.
Deeper layers represent increasingly abstract concepts and may be less reusable for a very different task.
Fix:
Consider earlier layers when the new dataset is substantially different.Confusing the convolutional base with the final prediction step.
The base transforms images into features; the new classifier interprets those features for the new classes.
Fix:
Keep the two jobs separate: feature production in the base and task-specific interpretation in the classifier.Using zero as a missing marker when zero is meaningful.
The network cannot reliably distinguish a real zero from missing data.
Fix:
Use a representation that separates missing values from meaningful zeros.
A strong reuse decision starts with three questions: Which part produces reusable visual representations? Which part is tied to the original classes? How similar is the new dataset to the original training data?
Guided practice
A new image dataset has different classes from the original dataset used to train a pretrained network. Describe the complete preparation and feature-extraction path: how the images become tensors, how their values are prepared, how missing values should be considered, which part of the pretrained network should be reused, and where the new classifier belongs.
Hints
- Begin with vectorization and tensor representation.
- Mention why small, relatively homogeneous values are useful.
- Check whether zero has a meaningful interpretation before using it for missing data.
- Separate the convolutional base from the original classifier.
- Explain how the new classifier uses the base output.
What do you think happens?
If the new image dataset differs substantially from the original training dataset, which part is generally the safer starting point for reuse: earlier convolutional layers or deeper convolutional layers?
Reveal answer
Answer: Earlier convolutional layers
Earlier layers tend to capture more general local patterns such as edges, colors, and textures. Deeper layers represent more abstract concepts and may be less reusable for a substantially different task.
Feature extraction in one pass
- Preprocessing changes raw data into numerical tensors that neural networks can use.
- Normalization keeps values small and helps avoid problems caused by incompatible feature ranges; scaling and standardizing describe different transformations.
- A missing-value marker is safe only when it cannot be confused with a meaningful data value such as zero.
- Feature extraction sends new images through a pretrained convolutional base and trains a new classifier on the base output.
- Earlier convolutional layers usually provide more general features, while deeper layers provide more abstract features whose usefulness depends more strongly on similarity between the old and new tasks.
Key Takeaways
- Neural-network inputs and targets must be converted into organized numerical tensors.
- Vectorization, normalization, and careful missing-value handling prepare raw data for learning.
- Scaling places values in a small range, whereas standardizing gives a feature a mean of 0 and a standard deviation of 1.
- A pretrained convolutional base supplies reusable visual representations, while a new classifier handles the new task's classes.
- Earlier layers are usually more general; the best reuse boundary depends on how different the new dataset is.