Image Classification with Deep Learning
A pretrained network is a saved network trained previously on a large dataset, such as ImageNet.
Starting with Learned Vision
Training an image classifier from the beginning requires the model to learn visual representations as well as the final class decisions. A pretrained network gives you a different starting point: it is a saved network that was previously trained on a large image dataset, such as ImageNet. Instead of discarding what that network learned, you can reuse the part that produces visual representations and train a new classifier for your own image classes.
Feature extraction separates two jobs: the convolutional base transforms images into learned visual features, while a new classifier interprets those features for the new classes.
Following an Image Through the Base
Feature extraction begins before the new classifier is trained. First, a new image is passed through the retained convolutional and pooling layers of the pretrained network. Those layers progressively transform the image into a learned visual representation. The output of the convolutional base then becomes the input to a separate classifier trained for the new image problem.
A New Two-Class Image Problem
Suppose a new task must distinguish between two image classes using a pretrained VGG16 network.
Select the base: Keep the VGG16 convolutional and pooling layers rather than its original densely connected classification section. In this feature-extraction setup, the selection is represented by include_top=False.
Process the new image: Pass the new image through the retained convolutional base. The base transforms the image into a learned visual representation.
Train for the new task: Use the base output as input to a new classifier trained for the two new classes.
The original visual feature-producing part is reused, while the classifier responsible for the original classes is replaced by a task-specific classifier.
Separating Base and Classifier
The boundary between reusable and task-specific parts is placed between the convolutional base and the original classifier. The convolutional base contains convolution and pooling layers that produce learned visual representations. The original classifier is tied to the classes from the first training task, so it should generally be replaced with a new classifier for the new class set.
Choosing a Layer Depth
Layer depth affects how general or task-specific a learned feature tends to be. Earlier convolutional layers usually capture local visual patterns such as edges, colors, and textures. Deeper layers capture increasingly abstract concepts, such as a cat ear or a dog eye. Because early features are more generic, they can be more reusable when the new dataset differs substantially from the original training data.
| Part of the network | Typical representation | Reuse implication |
|---|---|---|
| Earlier convolutional layers | Edges, colors, and textures | More generic and often more reusable |
| Deeper convolutional layers | More abstract concepts, such as a cat ear or dog eye | May be less reusable when the new task differs substantially |
| Original classifier | Representations tied to the original classes | Generally replaced for the new task |
Avoiding Reuse Errors
Keeping the original classifier unchanged for a new set of image classes.
The original classifier is tied to the classes from the first training task, not necessarily the classes in the new problem.
Fix:
Replace the original classifier with a new classifier trained using the output of the retained convolutional base.Including the original densely connected classification section when the goal is feature extraction.
The feature-extraction setup requires the new images to pass through the retained convolutional and pooling layers before a new classifier uses the resulting output.
Fix:
Use the feature-extraction boundary represented by include_top=False.Assuming the deepest available features are always the best features to reuse.
Deeper layers encode increasingly abstract concepts and may be less reusable for a substantially different task.
Fix:
Consider earlier convolutional layers when the new dataset differs substantially, because their features are more general.Describing feature extraction as if the new classifier processes the raw image directly.
The convolutional base transforms the image into learned visual features before the new classifier interprets them.
Fix:
Trace the image through the retained convolutional and pooling layers, then connect the base output to the new classifier.
Applying the Decision
What do you think happens?
A new image task uses classes that differ substantially from the classes in the dataset used to train the pretrained network. Which part is generally the safer starting point for reuse: an earlier convolutional layer or a deeper convolutional layer?
Reveal answer
Answer: An earlier convolutional layer
Earlier layers tend to capture more generic patterns such as edges, colors, and textures. Deeper layers represent more abstract concepts and may be less reusable when the new dataset differs substantially.
Explain the data path for a new image classification task in three parts: identify the reused component, describe the representation it produces, and state what component is trained for the new classes.
Hints
- Start with the convolutional base rather than the original classifier.
- Mention the convolutional and pooling layers.
- End with the new classifier that interprets the base output.
A new dataset is very different from the original training dataset. Decide whether earlier or deeper convolutional features are more appropriate as a starting point, and justify your choice using the kinds of patterns those layers represent.
Hints
- Compare generic local patterns with abstract concepts.
- Consider how similarity between the two datasets affects reuse.
Essential Takeaways
- A pretrained network is a saved network previously trained on a large dataset, such as ImageNet.
- Feature extraction sends new images through a retained convolutional base and uses the resulting visual representation as input to a new classifier.
- The convolutional base contains reusable convolutional and pooling layers, while the original classifier is tied to the original classes and is generally replaced.
- Earlier layers usually learn generic patterns such as edges, colors, and textures; deeper layers learn more abstract concepts.
- When the new dataset differs substantially from the original one, earlier convolutional features may provide a more appropriate starting point for reuse.
Key Takeaways
- A pretrained network provides learned visual representations from earlier training on a large image dataset.
- Feature extraction separates visual representation from task-specific classification.
- The convolutional base is reused, while a new classifier is trained for the new classes.
- Feature reuse depends on layer depth and on how similar the new dataset is to the original dataset.
- Earlier convolutional features are generally more generic, whereas deeper features are more abstract and task-related.