Concepts / Convolutional Neural Network Architecture

Convolutional Neural Network Architecture

A pretrained network is a saved network trained previously on a large dataset, such as ImageNet.

  • Programming

Starting with Learned Vision

Training an image classifier does not always begin with an untrained network. A pretrained network is a saved network that was trained previously on a large dataset, such as ImageNet. Instead of discarding what that network learned, you can reuse its convolutional base to produce visual representations for a new image problem.

The central idea is to separate visual representation from task-specific prediction. The pretrained convolutional base transforms images into learned visual features, while a new classifier interprets those features for the new set of classes.

Following an Image Through the Base

Feature extraction begins before the new classifier is trained. First, a new image is passed through the retained convolutional and pooling layers of the pretrained network. Those layers transform the image into a feature representation. The output of the convolutional base then becomes the input used to train a separate classifier for the new image problem.

passes throughcontinues throughproducesNew imageinputConvolutional layerslearned visual processingPooling layersretained baseFeaturerepresentationbase output
How does a new image move through successive convolutional and pooling layers, and what feature representation comes out at the end?

Tracing a New Image

A new image must be classified for a new image problem using a pretrained convolutional network. What happens during feature extraction?

Retain the convolutional base: Keep the part of the pretrained network made up of convolutional and pooling layers.

Pass in the new image: The new image moves through the retained base, which transforms it into a learned visual representation.

Use the base output: The resulting representation is passed to a separate classifier trained for the new set of classes.

The convolutional base performs visual feature extraction, and the new classifier performs task-specific interpretation.

Separating Features from Predictions

The convolutional base and the classifier have different jobs. The convolutional base contains convolution and pooling layers that produce learned visual representations. The original classifier is connected to the classes used during the first training task, so it is generally replaced by a new classifier for the new task.

entersfeature representationinterpretsNew imageConvolutional basereusable visual featuresNew classifiernew class interpretationNew-task prediction
Which layers contain general visual features, which layer makes the final task-specific prediction, and how are they connected?
Network partMain jobRole in a new task
Convolutional baseTransforms an image into learned visual featuresStarting point for reuse
Original classifierInterprets representations for the original classesGenerally replaced
New classifierInterprets the base output for the new classesTrained for the new image problem

Depth and Feature Reuse

Not every learned representation is equally reusable. Earlier convolutional layers usually capture more generic local visual patterns, such as edges, colors, and textures. Deeper layers capture increasingly abstract patterns, such as a cat ear or a dog eye. These deeper representations can still be useful, but they may be less suitable when the new dataset differs substantially from the original training data.

increasing abstractionincreasing specificityEarlier layersedges, colors, texturesHigher layersmore abstract patternsDeeper layersobject parts andtask-linked concepts
How do features change from edges and textures in early layers to object parts and task-specific patterns in deeper layers?

The best reuse boundary depends on how different the new dataset is from the original one. When the datasets are substantially different, the first few convolutional layers may be more appropriate because their features are more general. When the new problem is closer to the original visual domain, representations from higher layers may be more useful. The source material establishes the direction of this trade-off: depth brings abstraction, while earlier features are generally more reusable across differing image problems.

Choosing the Reuse Boundary

Selecting a pretrained network is not just a matter of keeping as many layers as possible. Begin with the convolutional base as the reuse candidate, remove the original classifier when setting up feature extraction, and attach a new classifier to the base output. Then consider the similarity between the original training data and the new dataset when deciding how far into the convolutional hierarchy to reuse representations.

original feature pathnew feature pathConvolutional baseretainedOriginal classifieroriginal classesConvolutional baseretainedNew classifiernew classes
What changes when the original classifier is removed and a new task-specific classifier is connected to the retained base?

Mistakes Beginners Make

  • Treating the entire pretrained network as the reusable component

    The original classifier is tied to the classes used during the first training task, while the new problem has a different set of classes.

    Fix: Use the convolutional base as the starting point for reuse and place a new classifier after its output.

  • Training the new classifier before extracting features from the base

    Feature extraction requires the base to transform each new image into a learned visual representation first.

    Fix: Trace the data path in order: new image, retained convolutional base, feature representation, new classifier.

  • Assuming the deepest features are always the best choice

    Deeper layers capture more abstract concepts and may be less reusable for a substantially different task.

    Fix: Consider earlier layers when the new dataset differs substantially, because their features are generally more generic.

  • Ignoring the feature-extraction boundary

    The feature-extraction configuration separates the convolutional base from the original classifier.

    Fix: For the VGG16 setup described here, use include_top=False so the original densely connected classification section is left out.

Check Your Architecture Reading

MEDIUM

A pretrained convolutional network is being adapted to a new image problem whose dataset differs substantially from the original training data. Identify the part that should be the starting point for reuse, the part that should generally be replaced, and the layer depth that may provide more reusable representations.

Hints
  • Separate the layers that produce visual features from the layer that interprets original classes.
  • Think about how abstraction changes as convolutional depth increases.
  • Use the relationship between dataset difference and feature generality.

A Complete Selection

Choose a reuse strategy for a new image dataset that differs substantially from the dataset used to train the pretrained network.

Choose the reusable component: Start with the pretrained convolutional base because it contains the convolutional and pooling layers that produce visual representations.

Replace the task-specific component: Leave out the original classifier and use a new classifier for the new set of classes.

Consider feature depth: Because the datasets differ substantially, earlier convolutional layers may be more appropriate because their learned patterns are more general.

Reuse begins with the convolutional base, uses a new classifier, and gives special consideration to earlier layers when the new visual problem differs substantially from the original one.

Key Takeaways

  1. A pretrained network is a saved network trained previously on a large dataset such as ImageNet.
  2. Feature extraction passes new images through the pretrained convolutional base before a new classifier is trained.
  3. The convolutional base produces reusable visual representations; the classifier interprets those representations for a particular set of classes.
  4. Earlier layers usually learn more generic patterns, while deeper layers learn more abstract concepts.
  5. When the new dataset differs substantially from the original one, earlier convolutional layers may provide more reusable features.

Key Takeaways

  • A pretrained network provides learned visual representations from earlier training on a large dataset.
  • During feature extraction, a new image travels through the retained convolutional and pooling layers, producing a feature representation.
  • The original classifier is generally replaced because it is tied to the original classes.
  • Earlier layers tend to be more generic, while deeper layers become more abstract and potentially less reusable for substantially different datasets.
  • A sound reuse decision considers both the convolutional-base boundary and the similarity between the old and new image datasets.