Densely Connected Layers and Convolution Layers
Convolution focuses on local patches rather than treating the entire image as one undivided pattern.
Try it: Convolution
How a convolution layer slides a small kernel over an image, multiplying and summing at each position to build a feature map.
How it works
- Place the 3×3 kernel over the top-left 3×3 window of the input.
- Multiply each input value by the kernel weight on top of it and add the nine products — that sum is one output cell.
- Slide the window by the stride and repeat until every position is covered.
- Zero padding adds a border so edge pixels are covered too; output size = ⌊(n + 2p − k) / s⌋ + 1.
Default run (11 steps): 5×5 input, 3×3 kernel, stride 1. Output is ⌊(5 + 2×0 − 3) / 1⌋ + 1 = 3 per side. … Done: the kernel visited all 9 positions and produced a 3×3 feature map.
Simplified: One 5×5 single-channel input and one 3×3 kernel. As in deep-learning libraries, the kernel is not flipped (strictly, cross-correlation).
Loading the simulation…
From Whole Images to Local Patches
A densely connected layer can learn global patterns involving all pixels of an image. A convolution layer takes a different approach: it examines small local regions, called patches, rather than treating the entire image as one enormous undivided object. This local focus lets the layer search for the same learned pattern at many positions.
The essential change in viewpoint is this: convolution asks what a local region contains, then repeats that question at different spatial locations.
Tracing a Filter Across a Feature Map
A convolution operation extracts a local patch from the input feature map. A learned transformation is applied to that patch, and the resulting response is recorded at the corresponding position in the output feature map. The operation then examines another patch and repeats the process.
One Transformation, Several Locations
Suppose a filter is examining a feature map. What happens as the filter moves from one local patch to another?
Select a patch: The convolution layer takes a small local region from the input feature map.
Apply the learned transformation: The filter evaluates that patch and produces a response for its location.
Move spatially: The layer selects another patch at a different position along the height or width of the feature map.
Reuse the transformation: The same filter transformation is applied again rather than learning a separate transformation for this new patch.
Build the output: The responses from the different locations form positions in an output response map.
The filter produces a spatial record of how strongly its learned pattern responds across the input.
Reading Feature-Map Axes
Convolutions operate over three-dimensional tensors called feature maps. Height and width describe spatial position. Depth, also called the channels axis, describes the feature channels carried by the map.
The output of a convolution layer is also a three-dimensional feature map. Its height and width continue to identify spatial positions. Its depth is determined by the layer and represents the filters and their responses, rather than specific RGB colors.
| Feature-map part | Meaning in an input or output map |
|---|---|
| Height | Vertical spatial position |
| Width | Horizontal spatial position |
| Depth | Channels |
| Output depth | Filters and their responses |
The roles of the three axes in convolutional feature maps.
Filters and Response Maps
A filter encodes a specific aspect of the input data. When a convolution layer uses several filters, each filter produces one channel in the output feature map. That channel is a response map: a two-dimensional grid showing how strongly the filter responds at different locations.
Interpreting Output Depth
A convolution layer uses several filters. What does each resulting output channel represent?
Start with one filter: One filter encodes one learned aspect of the input data.
Apply it across the input: The filter examines local patches at different spatial locations.
Collect its responses: The responses form a two-dimensional grid, called a response map.
Repeat for other filters: Each additional filter produces its own response map.
Stack the maps as output channels: The collection of response maps forms the depth of the output feature map.
Output depth represents filter responses. It does not simply carry the original image's RGB color channels.
Global Connections and Local Connections
The defining contrast is how much of the input an output unit considers. A densely connected layer can learn global patterns involving all pixels. A convolution layer learns patterns from local windows and applies the same local transformation across the input.
| Layer type | Pattern viewpoint | Transformation location |
|---|---|---|
| Densely connected layer | Global patterns involving all pixels | Connections use the entire input |
| Convolution layer | Patterns in local windows | The same transformation is reused across patches |
Translation and Pattern Hierarchies
Because the same local transformation is applied at different locations, convolutional networks can recognize a learned pattern when it appears in different parts of an image. A pattern learned in one corner can be recognized elsewhere without being learned as an entirely new pattern for every position. This is the translation-invariant behavior described for convolutional networks.
Convolutional networks also form spatial hierarchies of patterns. Earlier convolution layers can learn small local patterns such as edges. Later layers can combine features from earlier layers into larger patterns, and further layers can combine those into increasingly complex and abstract visual concepts. The progression is from small and local to larger and more abstract.
Common Misunderstandings
Thinking that a convolution layer examines the whole image as one undivided pattern.
Convolution focuses on small local windows or patches.
Fix:
Describe the operation as selecting local patches and applying a learned transformation across them.Assuming that a new transformation is learned for every spatial location.
The same transformation is reused across patches.
Fix:
Emphasize that the filter searches for the same learned pattern at multiple locations.Confusing depth with image height or width.
Height and width describe spatial position, while depth describes channels.
Fix:
Remember the three axes as height, width, and depth or channels.Interpreting every output channel as an RGB color channel.
Output channels represent filters and their responses.
Fix:
Interpret each output channel as a response map produced by a filter.Treating translation invariance and hierarchy as the same idea.
Translation invariance concerns recognizing patterns across locations, while hierarchy concerns combining features across layers.
Fix:
Keep location reuse and multi-layer feature combination as separate explanations.
Check Your Understanding
Explain the path from one input feature map to one output response map. Your explanation should mention a local patch, a filter, repeated spatial locations, and the output channel.
Hints
- Start by describing what the convolution layer selects from the input.
- Explain what is reused when the operation moves to another location.
- Finish by describing what the collection of responses represents.
A learner says, “The depth of an output feature map tells us the RGB color at each position.” Correct the statement and explain what output depth represents instead.
Hints
- Separate the meaning of input color channels from the meaning of output channels.
- Recall what each filter produces.
Describe how a network can move from detecting edges to recognizing increasingly complex visual concepts. Then explain why the same learned local pattern can be recognized in more than one image location.
Hints
- Use the progression from small and local to larger and more abstract.
- For the second part, focus on reuse of the same transformation across locations.
Key Takeaways
- Convolution examines local patches instead of treating an image as one undivided object.
- The same learned transformation is reused across patches at different height and width positions.
- A feature map has height, width, and depth or channels axes.
- Each filter produces a two-dimensional response map, and the collection of response maps forms the output depth.
- Convolutional networks recognize patterns across locations and build spatial hierarchies from small local features to larger, more abstract concepts.
Key Takeaways
- Convolution layers focus on local patches rather than analyzing an entire image as one global pattern.
- A filter applies the same learned transformation at many spatial locations.
- Height and width identify spatial position, while depth identifies channels.
- Each filter creates a response map, and output channels represent filter responses.
- Shared local transformations support translation-invariant recognition, while successive layers build spatial hierarchies from simple features to complex concepts.