Concepts / Pretrained ImageNet Models

Pretrained ImageNet Models

DeepDream changes an existing image by maximizing activations inside a convolutional network.

  • Programming

From Image to Dream

DeepDream does not begin by creating an image from nothing. It starts with an existing image and changes the image so that a convolutional network responds more strongly inside selected layers. As the pixels are repeatedly adjusted, patterns learned by the network become visually prominent. The result is an altered version of the starting image rather than an unrelated collection of new pixels.

The central idea is simple: use the network's internal activations as a reason to modify the image itself.

One Ascent Step

The network evaluates the current image and produces activations in its layers. DeepDream uses those activations to define an objective. Gradient ascent then determines how the input pixels should change so that the objective becomes larger. Repeating the adjustment makes the image a progressively stronger expression of features learned by the network.

network evaluatesobjective responsechanges pixelscontinue optimizationevaluate againExisting imagestarting pixelsLayer activationsmeasured responsesGradient ascentpixel adjustment directionModified imagestronger selected responsesRepeatmore ascent steps
How does the image change step by step as gradients are used to maximize the network's activations?

The important direction is from the objective back to the input pixels. DeepDream is not merely recording what the network sees. It uses the response to decide how to alter what the network will see next. Because the process is repeated, small pixel changes accumulate into visible distortions connected to patterns already available in the original image.

Whole Layers and Existing Images

AspectDeepDreamFilter visualization
Starting inputAn existing imageBlank, slightly noisy input
Optimization targetAn entire selected layerA specific filter in an upper layer
ResultA modified version of the starting imageAn image shaped around the selected filter response
Feature mixtureMany learned feature responses can contributeThe target is one specific filter

DeepDream and filter visualization use the same broad strategy: change the input in the direction that increases an internal network response. Their distinctive choices are different. Filter visualization can maximize one specific filter, while DeepDream maximizes an entire layer. Since a whole layer combines many learned feature responses, the resulting image can contain a mixture of visual patterns instead of one isolated feature.

Tracing a DeepDream Modification

Suppose an existing image is supplied to a pretrained convolutional network, and a selected layer is used as the DeepDream target. What happens during optimization?

Start: The original image is kept as the base input rather than being replaced by blank, slightly noisy input.

Measure: The network processes the image, and the activations in the selected layer contribute to the objective.

Ascend: Gradient ascent changes the input pixels in the direction that increases the objective.

Repeat: The modified image is evaluated again and adjusted again, so selected learned patterns become increasingly prominent.

The output is a distorted version of the original image whose visual changes are connected to patterns recognized by the selected network layer.

The Loss Behind the Change

Gradient ascent needs one numerical objective to increase. DeepDream defines this objective as a weighted sum of L2-based activation contributions from selected layers. Each layer has a coefficient, so its contribution can have more or less influence on the total objective. The cited formulation normalizes contributions by the number of activation values and leaves out border pixels when calculating the activation contribution, helping avoid border artifacts.

You can interpret the loss as a collection of layer measurements being combined into one score. First, activation values from a selected layer contribute an L2-based magnitude. That contribution is normalized over the activation values, with border positions excluded in the cited formulation. Then the layer's weight determines how strongly that contribution affects the total. Gradient ascent changes the image to make this combined score larger.

  • Selected layers provide the activation contributions.
  • Each contribution is based on the L2 magnitude of the layer's activations.
  • The contributions are normalized by the number of activation values.
  • Border activation positions are left out in the cited formulation.
  • Layer coefficients weight the contributions before they are combined.
  • The resulting weighted total is the objective increased by gradient ascent.

Why DeepDream Uses Octaves

DeepDream does not process the image at only one size. It defines several image scales called octaves. Processing begins with a smaller version and moves through increasingly larger versions. Each successive scale is 1.4 times the previous scale, which is 40 percent larger. Working across these scales allows feature maximization to act while the image is examined at different sizes, improving the quality of the visualization.

move to next scalemove to next scaleSmaller imageinitial octaveLarger image1.4 times previous scaleStill larger image1.4 times previous scale
How does the image move through different resolutions, and what changes between one octave and the next?

An octave is a processing scale, not a different image concept. The same DeepDream idea is applied while the image moves from a smaller resolution to successively larger resolutions.

Choosing the Network

DeepDream depends on the features learned by the convolutional network that evaluates the image. The source example uses the Inception V3 model available with Keras, while also noting that other pretrained convolutional networks can be used. Because different architectures learn different features, changing the pretrained network can change the appearance of the resulting visualization.

Common Reasoning Errors

  • Thinking DeepDream creates an image from blank input.

    DeepDream starts with an existing image. Blank, slightly noisy input is associated with the contrasting filter-visualization procedure described in the source.

    Fix: Remember that DeepDream repeatedly nudges an existing image toward stronger network responses.

  • Treating DeepDream as maximization of one filter.

    DeepDream maximizes an entire layer, which combines many learned feature responses.

    Fix: Use whole-layer maximization to explain why the output can contain a mixture of visual patterns.

  • Describing gradient ascent as changing the network weights.

    The source describes gradient ascent as repeatedly adjusting the input pixels in the direction that increases the objective.

    Fix: Track the changing object correctly: the image pixels are adjusted so the network's selected activations become larger.

  • Ignoring the scale progression.

    DeepDream defines several image scales called octaves and moves from smaller to increasingly larger versions.

    Fix: Include the octave sequence and the 1.4 scale factor between successive scales.

  • Treating the loss as one unweighted activation value.

    The objective is a weighted sum of L2-based activation contributions, with normalization over activation values and border positions omitted in the cited formulation.

    Fix: Explain how selected-layer contributions are measured and then combined according to their coefficients.

Check Your Model

MEDIUM

Explain, in your own words, why DeepDream is different from filter visualization even though both methods change an input to increase an internal network response. Include the starting input, the optimization target, and the role of gradient ascent.

Hints
  • Identify whether the input is an existing image or blank, slightly noisy input.
  • Compare a whole layer with a specific filter.
  • State what gradient ascent changes.
EASY

A DeepDream explanation mentions a single image size but says nothing about octaves. What important part of the method is missing, and why is it included?

Hints
  • Recall the name given to DeepDream's image scales.
  • State how successive scales change.
  • Connect multiple scales to visualization quality.

Key Takeaways

  1. DeepDream modifies an existing image by maximizing activations inside a convolutional network.
  2. Unlike single-filter visualization, it can maximize whole layers, allowing many learned feature responses to influence the result.
  3. Gradient ascent repeatedly changes input pixels in the direction that increases the objective.
  4. Octaves are successive image scales, beginning with a smaller version and increasing by a factor of 1.4 between scales.
  5. The objective combines weighted, normalized L2-based activation contributions from selected layers, with border positions omitted in the cited formulation.

Key Takeaways

  • DeepDream changes an existing image rather than starting from blank, slightly noisy input.
  • It increases responses from whole selected layers, so multiple learned features can shape the result.
  • Gradient ascent is the repeated pixel-adjustment process that raises the objective.
  • Octaves let the process work across increasingly larger image scales, with each scale 1.4 times the previous one.
  • The objective is a weighted combination of normalized L2-based activation contributions from selected layers.