Convolutional Neural Network Architecture
A pretrained network is a saved network trained previously on a large dataset, such as ImageNet.
Starting with Learned Vision
Training an image classifier does not always begin with an untrained network. A pretrained network is a saved network that was trained previously on a large dataset, such as ImageNet. Instead of discarding what that network learned, you can reuse its convolutional base to produce visual representations for a new image problem.
The central idea is to separate visual representation from task-specific prediction. The pretrained convolutional base transforms images into learned visual features, while a new classifier interprets those features for the new set of classes.
Following an Image Through the Base
Feature extraction begins before the new classifier is trained. First, a new image is passed through the retained convolutional and pooling layers of the pretrained network. Those layers transform the image into a feature representation. The output of the convolutional base then becomes the input used to train a separate classifier for the new image problem.
Tracing a New Image
A new image must be classified for a new image problem using a pretrained convolutional network. What happens during feature extraction?
Retain the convolutional base: Keep the part of the pretrained network made up of convolutional and pooling layers.
Pass in the new image: The new image moves through the retained base, which transforms it into a learned visual representation.
Use the base output: The resulting representation is passed to a separate classifier trained for the new set of classes.
The convolutional base performs visual feature extraction, and the new classifier performs task-specific interpretation.
Separating Features from Predictions
The convolutional base and the classifier have different jobs. The convolutional base contains convolution and pooling layers that produce learned visual representations. The original classifier is connected to the classes used during the first training task, so it is generally replaced by a new classifier for the new task.
| Network part | Main job | Role in a new task |
|---|---|---|
| Convolutional base | Transforms an image into learned visual features | Starting point for reuse |
| Original classifier | Interprets representations for the original classes | Generally replaced |
| New classifier | Interprets the base output for the new classes | Trained for the new image problem |
Depth and Feature Reuse
Not every learned representation is equally reusable. Earlier convolutional layers usually capture more generic local visual patterns, such as edges, colors, and textures. Deeper layers capture increasingly abstract patterns, such as a cat ear or a dog eye. These deeper representations can still be useful, but they may be less suitable when the new dataset differs substantially from the original training data.
The best reuse boundary depends on how different the new dataset is from the original one. When the datasets are substantially different, the first few convolutional layers may be more appropriate because their features are more general. When the new problem is closer to the original visual domain, representations from higher layers may be more useful. The source material establishes the direction of this trade-off: depth brings abstraction, while earlier features are generally more reusable across differing image problems.
Choosing the Reuse Boundary
Selecting a pretrained network is not just a matter of keeping as many layers as possible. Begin with the convolutional base as the reuse candidate, remove the original classifier when setting up feature extraction, and attach a new classifier to the base output. Then consider the similarity between the original training data and the new dataset when deciding how far into the convolutional hierarchy to reuse representations.
Mistakes Beginners Make
Treating the entire pretrained network as the reusable component
The original classifier is tied to the classes used during the first training task, while the new problem has a different set of classes.
Fix:
Use the convolutional base as the starting point for reuse and place a new classifier after its output.Training the new classifier before extracting features from the base
Feature extraction requires the base to transform each new image into a learned visual representation first.
Fix:
Trace the data path in order: new image, retained convolutional base, feature representation, new classifier.Assuming the deepest features are always the best choice
Deeper layers capture more abstract concepts and may be less reusable for a substantially different task.
Fix:
Consider earlier layers when the new dataset differs substantially, because their features are generally more generic.Ignoring the feature-extraction boundary
The feature-extraction configuration separates the convolutional base from the original classifier.
Fix:
For the VGG16 setup described here, use include_top=False so the original densely connected classification section is left out.
Check Your Architecture Reading
A pretrained convolutional network is being adapted to a new image problem whose dataset differs substantially from the original training data. Identify the part that should be the starting point for reuse, the part that should generally be replaced, and the layer depth that may provide more reusable representations.
Hints
- Separate the layers that produce visual features from the layer that interprets original classes.
- Think about how abstraction changes as convolutional depth increases.
- Use the relationship between dataset difference and feature generality.
A Complete Selection
Choose a reuse strategy for a new image dataset that differs substantially from the dataset used to train the pretrained network.
Choose the reusable component: Start with the pretrained convolutional base because it contains the convolutional and pooling layers that produce visual representations.
Replace the task-specific component: Leave out the original classifier and use a new classifier for the new set of classes.
Consider feature depth: Because the datasets differ substantially, earlier convolutional layers may be more appropriate because their learned patterns are more general.
Reuse begins with the convolutional base, uses a new classifier, and gives special consideration to earlier layers when the new visual problem differs substantially from the original one.
Key Takeaways
- A pretrained network is a saved network trained previously on a large dataset such as ImageNet.
- Feature extraction passes new images through the pretrained convolutional base before a new classifier is trained.
- The convolutional base produces reusable visual representations; the classifier interprets those representations for a particular set of classes.
- Earlier layers usually learn more generic patterns, while deeper layers learn more abstract concepts.
- When the new dataset differs substantially from the original one, earlier convolutional layers may provide more reusable features.
Key Takeaways
- A pretrained network provides learned visual representations from earlier training on a large dataset.
- During feature extraction, a new image travels through the retained convolutional and pooling layers, producing a feature representation.
- The original classifier is generally replaced because it is tied to the original classes.
- Earlier layers tend to be more generic, while deeper layers become more abstract and potentially less reusable for substantially different datasets.
- A sound reuse decision considers both the convolutional-base boundary and the similarity between the old and new image datasets.