Concepts / Densely Connected Layers and Convolution Layers

Densely Connected Layers and Convolution Layers

Convolution focuses on local patches rather than treating the entire image as one undivided pattern.

  • Programming
Interactive lab

Try it: Convolution

How a convolution layer slides a small kernel over an image, multiplying and summing at each position to build a feature map.

How it works

  1. Place the 3×3 kernel over the top-left 3×3 window of the input.
  2. Multiply each input value by the kernel weight on top of it and add the nine products — that sum is one output cell.
  3. Slide the window by the stride and repeat until every position is covered.
  4. Zero padding adds a border so edge pixels are covered too; output size = ⌊(n + 2p − k) / s⌋ + 1.

Default run (11 steps): 5×5 input, 3×3 kernel, stride 1. Output is ⌊(5 + 2×0 − 3) / 1⌋ + 1 = 3 per side. … Done: the kernel visited all 9 positions and produced a 3×3 feature map.

Simplified: One 5×5 single-channel input and one 3×3 kernel. As in deep-learning libraries, the kernel is not flipped (strictly, cross-correlation).

Educational simulation

Loading the simulation…

From Whole Images to Local Patches

A densely connected layer can learn global patterns involving all pixels of an image. A convolution layer takes a different approach: it examines small local regions, called patches, rather than treating the entire image as one enormous undivided object. This local focus lets the layer search for the same learned pattern at many positions.

select local regionmove across height or widthmove againrecord responseInput feature mapheight × width × depthPatch 1local regionOutput positionsresponsesPatch 2next locationPatch 3another location
What patch of the input feature map is selected next as a filter moves across height and width?

The essential change in viewpoint is this: convolution asks what a local region contains, then repeats that question at different spatial locations.

Tracing a Filter Across a Feature Map

A convolution operation extracts a local patch from the input feature map. A learned transformation is applied to that patch, and the resulting response is recorded at the corresponding position in the output feature map. The operation then examines another patch and repeats the process.

One Transformation, Several Locations

Suppose a filter is examining a feature map. What happens as the filter moves from one local patch to another?

Select a patch: The convolution layer takes a small local region from the input feature map.

Apply the learned transformation: The filter evaluates that patch and produces a response for its location.

Move spatially: The layer selects another patch at a different position along the height or width of the feature map.

Reuse the transformation: The same filter transformation is applied again rather than learning a separate transformation for this new patch.

Build the output: The responses from the different locations form positions in an output response map.

The filter produces a spatial record of how strongly its learned pattern responds across the input.

applyproduceapplyproducePatch Alocation oneFiltersame transformationResponse Aone positionPatch Blocation twoFiltersame transformationResponse Banother position
How is the same filter operation reused at different spatial locations instead of learning a separate transformation for every patch?

Reading Feature-Map Axes

Convolutions operate over three-dimensional tensors called feature maps. Height and width describe spatial position. Depth, also called the channels axis, describes the feature channels carried by the map.

axisaxisaxisHeightvertical positionFeature mapthree-dimensional tensorWidthhorizontal positionDepthchannels
Which dimension represents vertical position, horizontal position, and feature channels in a feature map?

The output of a convolution layer is also a three-dimensional feature map. Its height and width continue to identify spatial positions. Its depth is determined by the layer and represents the filters and their responses, rather than specific RGB colors.

Feature-map partMeaning in an input or output map
HeightVertical spatial position
WidthHorizontal spatial position
DepthChannels
Output depthFilters and their responses

The roles of the three axes in convolutional feature maps.

Filters and Response Maps

A filter encodes a specific aspect of the input data. When a convolution layer uses several filters, each filter produces one channel in the output feature map. That channel is a response map: a two-dimensional grid showing how strongly the filter responds at different locations.

apply across patchesapply across patchesproduceproduceInput feature mapmany local patchesFilter 1learned aspectResponse map 1one output channelFilter 2learned aspectResponse map 2one output channel
How does each filter transform many input patches into the corresponding positions of one output response map?

Interpreting Output Depth

A convolution layer uses several filters. What does each resulting output channel represent?

Start with one filter: One filter encodes one learned aspect of the input data.

Apply it across the input: The filter examines local patches at different spatial locations.

Collect its responses: The responses form a two-dimensional grid, called a response map.

Repeat for other filters: Each additional filter produces its own response map.

Stack the maps as output channels: The collection of response maps forms the depth of the output feature map.

Output depth represents filter responses. It does not simply carry the original image's RGB color channels.

Global Connections and Local Connections

The defining contrast is how much of the input an output unit considers. A densely connected layer can learn global patterns involving all pixels. A convolution layer learns patterns from local windows and applies the same local transformation across the input.

global connectionlocal connectionEntire inputall pixelsOutput unitglobal patternLocal patchsmall regionOutput positionlocal response
What is the difference between connecting every input value to every output unit and connecting each output to only a local patch?
Layer typePattern viewpointTransformation location
Densely connected layerGlobal patterns involving all pixelsConnections use the entire input
Convolution layerPatterns in local windowsThe same transformation is reused across patches

Translation and Pattern Hierarchies

Because the same local transformation is applied at different locations, convolutional networks can recognize a learned pattern when it appears in different parts of an image. A pattern learned in one corner can be recognized elsewhere without being learned as an entirely new pattern for every position. This is the translation-invariant behavior described for convolutional networks.

combinecombine furtherdetect across positionsEdgessmall local patternsLarger patternscombined featuresVisual conceptsabstract patternsDifferent locationsshared detection
How can lower-level features detected in different locations combine into larger patterns while preserving recognition when an object shifts position?

Convolutional networks also form spatial hierarchies of patterns. Earlier convolution layers can learn small local patterns such as edges. Later layers can combine features from earlier layers into larger patterns, and further layers can combine those into increasingly complex and abstract visual concepts. The progression is from small and local to larger and more abstract.

Common Misunderstandings

  • Thinking that a convolution layer examines the whole image as one undivided pattern.

    Convolution focuses on small local windows or patches.

    Fix: Describe the operation as selecting local patches and applying a learned transformation across them.

  • Assuming that a new transformation is learned for every spatial location.

    The same transformation is reused across patches.

    Fix: Emphasize that the filter searches for the same learned pattern at multiple locations.

  • Confusing depth with image height or width.

    Height and width describe spatial position, while depth describes channels.

    Fix: Remember the three axes as height, width, and depth or channels.

  • Interpreting every output channel as an RGB color channel.

    Output channels represent filters and their responses.

    Fix: Interpret each output channel as a response map produced by a filter.

  • Treating translation invariance and hierarchy as the same idea.

    Translation invariance concerns recognizing patterns across locations, while hierarchy concerns combining features across layers.

    Fix: Keep location reuse and multi-layer feature combination as separate explanations.

Check Your Understanding

EASY

Explain the path from one input feature map to one output response map. Your explanation should mention a local patch, a filter, repeated spatial locations, and the output channel.

Hints
  • Start by describing what the convolution layer selects from the input.
  • Explain what is reused when the operation moves to another location.
  • Finish by describing what the collection of responses represents.
MEDIUM

A learner says, “The depth of an output feature map tells us the RGB color at each position.” Correct the statement and explain what output depth represents instead.

Hints
  • Separate the meaning of input color channels from the meaning of output channels.
  • Recall what each filter produces.
MEDIUM

Describe how a network can move from detecting edges to recognizing increasingly complex visual concepts. Then explain why the same learned local pattern can be recognized in more than one image location.

Hints
  • Use the progression from small and local to larger and more abstract.
  • For the second part, focus on reuse of the same transformation across locations.

Key Takeaways

  1. Convolution examines local patches instead of treating an image as one undivided object.
  2. The same learned transformation is reused across patches at different height and width positions.
  3. A feature map has height, width, and depth or channels axes.
  4. Each filter produces a two-dimensional response map, and the collection of response maps forms the output depth.
  5. Convolutional networks recognize patterns across locations and build spatial hierarchies from small local features to larger, more abstract concepts.

Key Takeaways

  • Convolution layers focus on local patches rather than analyzing an entire image as one global pattern.
  • A filter applies the same learned transformation at many spatial locations.
  • Height and width identify spatial position, while depth identifies channels.
  • Each filter creates a response map, and output channels represent filter responses.
  • Shared local transformations support translation-invariant recognition, while successive layers build spatial hierarchies from simple features to complex concepts.