Activation Functions in Deep Neural Networks
DeepDream changes an existing image by maximizing activations inside a convolutional network.
From Image to Dream
DeepDream does not begin by drawing a new image from nothing. It starts with an existing image, sends that image through a convolutional neural network, and changes the image so that selected internal activations become larger. The result is an altered version of the original in which patterns recognized by the network become visually prominent.
The central direction of the process is from the network back to the pixels: the network evaluates the image, and the image is adjusted to produce stronger internal responses.
The Optimization Target
DeepDream needs a single numerical objective that tells it what to increase. That objective is formed from activation magnitudes in selected layers. Each selected layer contributes a magnitude based on its activations, and a coefficient determines how strongly that layer contributes. The contributions are normalized by the number of activation values. In the cited formulation, border activation positions are excluded so that border effects do not contribute to the activation calculation.
Two Layers, One Objective
Imagine that DeepDream has selected two layers. The first layer contributes an activation magnitude with one coefficient, and the second contributes another activation magnitude with its own coefficient.
Collect: The method gathers activation values from both selected layers.
Measure: It forms an L2-based magnitude contribution for each layer, using the nonborder activation positions in the cited formulation.
Weight: Each layer contribution is scaled by its coefficient, so the layers do not have to influence the objective equally.
Combine: The weighted contributions are combined into one objective. That single value is the target increased by gradient ascent.
The objective represents several selected layers at once rather than a single isolated filter response.
Pixel Updates by Gradient Ascent
After defining the objective, DeepDream determines how changes to the input image would change that objective. Gradient ascent then adjusts the image pixels in the direction that increases the objective. This update is repeated, so the original image progressively becomes a stronger expression of the patterns represented by the selected layer or layers.
What do you think happens?
Suppose the image is updated repeatedly in the direction that increases the selected activations. What should happen to patterns recognized by those activations?
Reveal answer
Answer: Those patterns should become increasingly visually prominent in the altered image.
Each update is chosen to increase the objective, and the objective is built from selected activation magnitudes. Repeating the updates therefore turns the original image into a stronger expression of the selected learned features.
Whole Layers Versus Single Filters
| Method | Starting input | Target | Visual result |
|---|---|---|---|
| DeepDream | An existing image | An entire selected layer | The starting image is altered toward a mixture of learned visual patterns |
| Filter visualization | Blank, slightly noisy input | A specific filter | The input is synthesized to reveal one isolated feature response |
Both methods use the broad idea of changing an input to increase an internal response. Their important difference is the target and the starting point. DeepDream retains the original image and maximizes a whole layer, whose many learned feature responses can produce a mixture of patterns. Filter visualization instead begins with blank, slightly noisy input and can maximize one specific filter.
The Octave-by-Octave Process
DeepDream does not process the image at only one size. It defines several image scales called octaves. Processing begins with a smaller version and continues through increasingly larger versions. Each successive scale is 1.4 times the previous scale, which is 40 percent larger. At each scale, feature maximization can act while the image is examined at that size. Working across these scales improves the quality of the visualization by allowing the process to act on the image at different sizes.
Following an Image Through Octaves
Trace what happens when an existing image is processed through several increasingly larger scales.
Begin smaller: DeepDream first works with a smaller version of the existing image.
Maximize activations: Gradient ascent changes the image at that scale so that the selected activations become stronger.
Increase scale: The process moves to a larger image scale. Each successive scale is 1.4 times the previous one.
Continue: The image is again processed at the new scale, allowing feature maximization to act while the image is examined at another size.
The final visualization reflects repeated activation maximization across multiple image sizes rather than one pass at one resolution.
Forward Evaluation and Feedback
The network remains the evaluator, while the image becomes the object being optimized. A forward evaluation produces activations inside the convolutional network. The selected activations contribute to the objective. Gradient ascent uses how that objective changes with respect to the image pixels to decide how the pixels should be adjusted. The network is then evaluated again on the updated image, creating a repeated feedback process.
Mistakes in Interpretation
Treating DeepDream as an image classifier that leaves the input unchanged.
DeepDream changes the pixels so that selected internal activations become larger.
Fix:
Think of the network as an evaluator whose responses define a target for modifying the image.Calling DeepDream single-filter visualization.
DeepDream maximizes an entire layer, combining many learned feature responses.
Fix:
Identify both distinctions: DeepDream starts from an existing image and targets a whole layer.Assuming DeepDream starts from blank, slightly noisy input.
That starting point describes the filter-visualization contrast, whereas DeepDream retains an existing image.
Fix:
Track the original image as the base that is repeatedly nudged toward stronger responses.Thinking gradient ascent changes network weights.
The described process repeatedly adjusts the input pixels in the direction that increases the objective.
Fix:
Separate the fixed evaluation role of the pretrained network from the optimization of the image pixels.Ignoring image scale.
DeepDream uses several scales called octaves, with each successive scale 1.4 times the previous scale.
Fix:
Include the progression from a smaller version through increasingly larger versions.
Practice the Mechanism
Explain, in order, what happens when DeepDream receives an existing image and is asked to maximize activations from a selected layer. Include the objective, gradient ascent, and octaves in your explanation.
Hints
- Start with the existing image rather than noise.
- Mention that selected layer activations contribute to a weighted objective.
- Explain that gradients determine how to adjust the image pixels.
- End by describing processing across increasingly larger scales.
Compare DeepDream with filter visualization using two categories: starting input and optimization target. Then explain why DeepDream can produce a mixture of visual patterns.
Hints
- DeepDream retains an existing image.
- Filter visualization begins with blank, slightly noisy input.
- DeepDream maximizes a whole layer, while filter visualization can maximize one specific filter.
- A whole layer combines many learned feature responses.
Key Takeaways
- DeepDream modifies an existing image so that selected activations inside a convolutional network become larger.
- Its objective is a weighted combination of L2-based activation contributions from selected layers, with normalization by activation count and border positions excluded in the cited formulation.
- Gradient ascent repeatedly changes the input pixels in the direction that increases the objective.
- DeepDream targets whole layers, unlike filter visualization, which can target one filter and begin from blank, slightly noisy input.
- Octaves are progressively larger image scales; each successive scale is 1.4 times the previous scale, allowing feature maximization at multiple sizes.
Key Takeaways
- DeepDream turns an existing image into a stronger expression of patterns learned by a convolutional neural network.
- It maximizes whole-layer activations rather than only one filter response.
- Gradient ascent uses the objective's response to pixel changes to repeatedly update the image.
- The objective combines weighted activation magnitudes from selected layers.
- Processing through octaves exposes the image to feature maximization at several progressively larger scales.