Concepts / Image Classification with Deep Learning

Image Classification with Deep Learning

A pretrained network is a saved network trained previously on a large dataset, such as ImageNet.

  • Programming

Starting with Learned Vision

Training an image classifier from the beginning requires the model to learn visual representations as well as the final class decisions. A pretrained network gives you a different starting point: it is a saved network that was previously trained on a large image dataset, such as ImageNet. Instead of discarding what that network learned, you can reuse the part that produces visual representations and train a new classifier for your own image classes.

Feature extraction separates two jobs: the convolutional base transforms images into learned visual features, while a new classifier interprets those features for the new classes.

trainslearnsusesfeedsLarge image datasetImageNetPretrained networksaved networkVisual representationlearned featuresNew imagesnew taskNew classifiernew classes
How can representations learned from a large dataset help classify images in a new task?

Following an Image Through the Base

Feature extraction begins before the new classifier is trained. First, a new image is passed through the retained convolutional and pooling layers of the pretrained network. Those layers progressively transform the image into a learned visual representation. The output of the convolutional base then becomes the input to a separate classifier trained for the new image problem.

enterspasses throughpasses throughproducesNew imageinputEarly convolutionvisual patternsPooling layersretained baseDeeper convolutionabstract patternsBase outputlearned features
How does a new image move through successive convolutional layers, and what feature representation comes out at the end?

A New Two-Class Image Problem

Suppose a new task must distinguish between two image classes using a pretrained VGG16 network.

Select the base: Keep the VGG16 convolutional and pooling layers rather than its original densely connected classification section. In this feature-extraction setup, the selection is represented by include_top=False.

Process the new image: Pass the new image through the retained convolutional base. The base transforms the image into a learned visual representation.

Train for the new task: Use the base output as input to a new classifier trained for the two new classes.

The original visual feature-producing part is reused, while the classifier responsible for the original classes is replaced by a task-specific classifier.

Separating Base and Classifier

The boundary between reusable and task-specific parts is placed between the convolutional base and the original classifier. The convolutional base contains convolution and pooling layers that produce learned visual representations. The original classifier is tied to the classes from the first training task, so it should generally be replaced with a new classifier for the new class set.

entersproducesfeedswas connected toNew imageinputConvolutional basereusedBase outputvisual featuresNew classifiernew classesOriginal classifieroriginal classes
Which parts of a pretrained network are reused, which part is replaced or retrained, and how are they connected?

Choosing a Layer Depth

Layer depth affects how general or task-specific a learned feature tends to be. Earlier convolutional layers usually capture local visual patterns such as edges, colors, and textures. Deeper layers capture increasingly abstract concepts, such as a cat ear or a dog eye. Because early features are more generic, they can be more reusable when the new dataset differs substantially from the original training data.

becomes more abstractbecomes more task-relatedoften supportsdepends on similarityEarly layersedges, colors, texturesHigher layersmore abstract patternsDeeper layerscat ear, dog eyeFeature reusedepends on datasetsimilarity
What changes from early to late convolutional layers, and why are early features generally more reusable across tasks?
Part of the networkTypical representationReuse implication
Earlier convolutional layersEdges, colors, and texturesMore generic and often more reusable
Deeper convolutional layersMore abstract concepts, such as a cat ear or dog eyeMay be less reusable when the new task differs substantially
Original classifierRepresentations tied to the original classesGenerally replaced for the new task

Avoiding Reuse Errors

  • Keeping the original classifier unchanged for a new set of image classes.

    The original classifier is tied to the classes from the first training task, not necessarily the classes in the new problem.

    Fix: Replace the original classifier with a new classifier trained using the output of the retained convolutional base.

  • Including the original densely connected classification section when the goal is feature extraction.

    The feature-extraction setup requires the new images to pass through the retained convolutional and pooling layers before a new classifier uses the resulting output.

    Fix: Use the feature-extraction boundary represented by include_top=False.

  • Assuming the deepest available features are always the best features to reuse.

    Deeper layers encode increasingly abstract concepts and may be less reusable for a substantially different task.

    Fix: Consider earlier convolutional layers when the new dataset differs substantially, because their features are more general.

  • Describing feature extraction as if the new classifier processes the raw image directly.

    The convolutional base transforms the image into learned visual features before the new classifier interprets them.

    Fix: Trace the image through the retained convolutional and pooling layers, then connect the base output to the new classifier.

feedsends withConvolutional baseretainedNew classifiernew classesFull pretrainednetworkoriginal top includedOriginal classifieroriginal classes
What can go wrong if the wrong layers are selected or the original classifier is reused unchanged?

Applying the Decision

What do you think happens?

A new image task uses classes that differ substantially from the classes in the dataset used to train the pretrained network. Which part is generally the safer starting point for reuse: an earlier convolutional layer or a deeper convolutional layer?

  • An earlier convolutional layer
  • A deeper convolutional layer
  • The original classifier
Reveal answer

Answer: An earlier convolutional layer

Earlier layers tend to capture more generic patterns such as edges, colors, and textures. Deeper layers represent more abstract concepts and may be less reusable when the new dataset differs substantially.

EASY

Explain the data path for a new image classification task in three parts: identify the reused component, describe the representation it produces, and state what component is trained for the new classes.

Hints
  • Start with the convolutional base rather than the original classifier.
  • Mention the convolutional and pooling layers.
  • End with the new classifier that interprets the base output.
MEDIUM

A new dataset is very different from the original training dataset. Decide whether earlier or deeper convolutional features are more appropriate as a starting point, and justify your choice using the kinds of patterns those layers represent.

Hints
  • Compare generic local patterns with abstract concepts.
  • Consider how similarity between the two datasets affects reuse.

Essential Takeaways

  1. A pretrained network is a saved network previously trained on a large dataset, such as ImageNet.
  2. Feature extraction sends new images through a retained convolutional base and uses the resulting visual representation as input to a new classifier.
  3. The convolutional base contains reusable convolutional and pooling layers, while the original classifier is tied to the original classes and is generally replaced.
  4. Earlier layers usually learn generic patterns such as edges, colors, and textures; deeper layers learn more abstract concepts.
  5. When the new dataset differs substantially from the original one, earlier convolutional features may provide a more appropriate starting point for reuse.

Key Takeaways

  • A pretrained network provides learned visual representations from earlier training on a large image dataset.
  • Feature extraction separates visual representation from task-specific classification.
  • The convolutional base is reused, while a new classifier is trained for the new classes.
  • Feature reuse depends on layer depth and on how similar the new dataset is to the original dataset.
  • Earlier convolutional features are generally more generic, whereas deeper features are more abstract and task-related.