Fine-Tuning a Pretrained Convnet
A pretrained network is a saved network trained previously on a large dataset, such as ImageNet.
Starting with Learned Vision
Training an image model from an uninitialized state is not the only way to solve a new image problem. A pretrained network is a saved network that was trained previously on a large dataset, such as ImageNet. Its convolutional layers have already learned visual representations, so a new task can begin with those learned representations instead of starting with an entirely uninitialized model.
The useful division is between representation and interpretation. The convolutional base transforms an image into learned visual features. A classifier then interprets those features for the new set of classes. This separation is especially useful when the new dataset is smaller than the dataset used for the original training.
Reuse does not mean that every original part of the network remains appropriate. The convolutional base is the starting point for reuse, while the original classifier is tied to the classes from the first task and should generally be replaced for a new classification problem.
Following an Image Through the Base
During feature extraction, a new image is passed through the retained convolutional and pooling layers. The base changes the image into a learned visual representation. A separate classifier receives that representation and learns how to interpret it for the new image problem.
Tracing a New Image
Suppose a pretrained convnet is being reused for a new set of image classes.
Provide the image: The new image enters the pretrained convolutional base.
Transform the image: The convolutional and pooling layers produce learned visual features.
Interpret the features: A new classifier receives the base output and learns to associate it with the classes in the new problem.
The base performs visual representation, while the new classifier performs task-specific interpretation.
Separating Base and Classifier
Feature extraction is the setup in which new images pass through the pretrained convolutional base and a new classifier is trained on the base output. The original densely connected classification section is excluded from the reused VGG16 base when the configuration uses include_top=False.
| Part | Role in the new task | Typical treatment in feature extraction |
|---|---|---|
| Convolutional base | Transforms images into learned visual features | Selected for reuse |
| Original classifier | Interprets features for the original classes | Excluded or replaced |
| New classifier | Interprets features for the target classes | Trained for the new problem |
The boundary between reusable representation and task-specific classification
Reading Feature Depth
Layer depth affects reuse because the kind of representation changes through the convolutional base. Earlier convolutional layers usually capture generic local patterns such as edges, colors, and textures. Deeper layers capture more abstract concepts, such as a cat ear or a dog eye. Generic patterns are more likely to remain useful when the new dataset differs from the original dataset.
Imagine reusing a convnet for a target dataset that differs substantially from the original training data. A reuse boundary nearer the first few convolutional layers may be more appropriate because those layers represent more general visual patterns. Higher layers may encode increasingly abstract concepts that are less reusable for a very different task.
Do not treat the deepest available feature as automatically the best feature for every target dataset. The appropriate reuse boundary depends partly on how different the new dataset is from the original one.
Adapting the Pretrained State
Fine-tuning is a technique for adapting a pretrained convolutional neural network to a target task. It begins with a convnet that has already been trained and uses that pretrained model as the starting point for further adjustment rather than treating the model as entirely new.
The adaptation can be described as a state transition. The starting state is a pretrained convnet. During the target-task stage, some parts may retain their existing parameters while other parts are updated. The exact retained and updated parts depend on the fine-tuning procedure; there is no single fixed layer list that defines every fine-tuning setup.
State Trace for Fine-Tuning
Describe how a pretrained convnet changes while it is adapted to a target image task.
Pretrained state: The model begins with parameters learned from an earlier large image dataset.
Task-specific preparation: The original classification section is not treated as the classifier for the new class set; a task-specific classifier is used instead.
Selective adjustment: Some model parts may retain their existing parameters while other parts are updated for the target task.
Adapted state: The result is still based on the pretrained convnet, but it has been adjusted toward the target task.
Fine-tuning means adapting a pretrained starting point, not rebuilding the entire convnet from an uninitialized state.
Mistakes in Reuse Decisions
Treating the original classifier as the reusable answer for the new task.
The original classifier is tied to the classes used during the first training task.
Fix:
Reuse the convolutional base as the starting visual representation and place a new classifier after its output.Assuming the deepest convolutional features are always the most reusable.
Deeper layers represent increasingly abstract concepts and may be less reusable for a substantially different task.
Fix:
Consider earlier layers when the target dataset differs substantially, because their patterns are generally more generic.Describing fine-tuning as updating every parameter in every situation.
Fine-tuning procedures can retain some parts and update others.
Fix:
State which parts are retained and which are updated for the particular procedure being discussed.Confusing feature extraction with fine-tuning.
Feature extraction uses the base to produce representations and trains a new classifier, while fine-tuning refers more broadly to adapting a pretrained convnet with selected parts potentially updated.
Fix:
Name the model state and identify whether the convolutional base is being retained or whether selected parts are also being adjusted.
Practice the Distinction
A new image dataset has classes different from those used to train a saved convnet. Explain which part you would begin by reusing, which part you would generally replace, and why the best reuse boundary might move toward earlier convolutional layers if the datasets differ substantially.
Hints
- Separate visual representation from class interpretation.
- Remember that the original classifier is tied to the original classes.
- Compare the generality of early features with the abstraction of deeper features.
Describe a fine-tuning state trace in four stages: the pretrained starting point, the task-specific classifier decision, the parts retained during adaptation, and the parts updated for the target task.
Hints
- Fine-tuning starts from a trained convnet.
- Do not assume that every layer must be updated.
- Make the retained-versus-updated distinction explicit.
Key Takeaways
- A pretrained network is a saved network trained previously on a large dataset, and its learned visual representations can provide a starting point for a new task.
- The convolutional base transforms a new image into learned features, while a task-specific classifier interprets those features for the new classes.
- Earlier layers usually capture more generic patterns; deeper layers capture more abstract concepts and may be less reusable when datasets differ substantially.
- Feature extraction trains a new classifier on the output of the pretrained base, while fine-tuning adapts a pretrained convnet with some parts retained and other parts potentially updated.
- A precise explanation of fine-tuning must identify the procedure's retained and updated parts rather than assuming one universal layer policy.
Key Takeaways
- A pretrained convnet supplies learned visual representations that can be reused for a new image task.
- Feature extraction sends new images through the convolutional base and trains a new classifier on the resulting features.
- The original classifier is tied to the original classes and is generally replaced for a new class set.
- Early convolutional features are usually more generic, while deeper features are more abstract and potentially less reusable across substantially different datasets.
- Fine-tuning adapts a pretrained starting point by retaining some parts and updating others according to the specific procedure.