Concepts / Feature Maps and Filters

Feature Maps and Filters

Convolution focuses on local patches rather than treating the entire image as one undivided pattern.

  • Programming
Interactive lab

Try it: Convolution

How a convolution layer slides a small kernel over an image, multiplying and summing at each position to build a feature map.

How it works

  1. Place the 3×3 kernel over the top-left 3×3 window of the input.
  2. Multiply each input value by the kernel weight on top of it and add the nine products — that sum is one output cell.
  3. Slide the window by the stride and repeat until every position is covered.
  4. Zero padding adds a border so edge pixels are covered too; output size = ⌊(n + 2p − k) / s⌋ + 1.

Default run (11 steps): 5×5 input, 3×3 kernel, stride 1. Output is ⌊(5 + 2×0 − 3) / 1⌋ + 1 = 3 per side. … Done: the kernel visited all 9 positions and produced a 3×3 feature map.

Simplified: One 5×5 single-channel input and one 3×3 kernel. As in deep-learning libraries, the kernel is not flipped (strictly, cross-correlation).

Educational simulation

Loading the simulation…

From Whole Images to Local Patches

A convolution layer does not treat an image as one enormous, undivided object. Instead, it examines small local regions called patches. For an image, each patch is a small two-dimensional window of the input. The layer applies a learned transformation to these local regions and records the resulting responses in an output feature map.

The central idea is local examination plus reuse: convolution looks at one patch at a time, but uses the same learned transformation at many locations.

Following the Moving Patch

Imagine an input feature map laid out as a spatial field. A convolution operation selects a local patch at one position. It applies the learned transformation to that patch and places the resulting response at the corresponding position in the output. The operation then examines another patch at a different location and repeats the same process. As this continues across the input, the output records where the transformation responds strongly or weakly.

selectmovetransformtransformInput feature mapspatial fieldPatch Alocal windowOutput position AresponsePatch Bnext local windowOutput position Bresponse
Which local patch is examined at each step, and where does the operation continue next?

Tracing Two Patch Positions

A convolution layer examines one local patch and then another patch at a different location. What does it reuse, and what changes?

Select the first patch: The layer examines a small local region of the input feature map.

Apply the learned transformation: The layer applies its learned transformation to that patch and records a response at the corresponding output position.

Select another patch: The examined window moves to a different location in the input.

Reuse the transformation: The same transformation is applied again rather than being redesigned for the new patch.

The patch location changes, but the learned transformation is reused. The output therefore records responses at multiple spatial positions.

Reading Feature Map Axes

Convolutions operate over three-dimensional tensors called feature maps. Height and width describe spatial position: they locate a response or feature across the vertical and horizontal dimensions. The third axis is depth, also called the channels axis. Depth represents channels rather than another spatial direction.

contains axiscontains axiscontains axisFeature mapthree-dimensional tensorHeightvertical positionWidthhorizontal positionDepthchannels
What do height, width, and depth contain, and how are they arranged in a feature map?
AxisMeaning in a feature map
HeightSpatial position along one direction
WidthSpatial position along the other direction
DepthChannels containing feature information

Filters and Response Maps

A filter encodes a specific aspect of the input data. When a convolution layer uses several filters, each filter produces a channel in the output. That channel is a response map: a two-dimensional grid showing the response of that filter at different locations in the input.

The output feature map keeps three dimensions. Its height and width identify spatial positions, while its depth is chosen by the layer and contains the filter responses. One output channel can therefore be read as a map of where one learned filter finds its preferred aspect of the input.

examines patchesproducesarrange spatiallyInput feature maplocal patchesFilterlearned aspectFilter responsesone value per positionResponse mapoutput channel
How does one filter turn responses from input patches into a channel of the output feature map?

Interpreting One Output Channel

A convolution layer uses several filters. How should one resulting output channel be interpreted?

Identify the filter: Choose one filter that encodes a particular aspect of the input data.

Read its local responses: At different input locations, the filter produces responses to the patches it examines.

Arrange the responses: Those responses form a two-dimensional grid aligned with spatial positions.

The selected output channel is a response map showing where the chosen filter responds across the input.

Reusing One Transformation

The transformation is not redesigned for every patch. It is reused across the input, allowing the same learned pattern to be sought in multiple places. This reuse is what lets one filter create a response map instead of learning a separate transformation for every spatial position.

applyapplyapplyproduceproduceproducePatch oneinput locationSame filterreused transformationResponse oneoutput positionPatch twoinput locationResponse twooutput positionPatch threeinput locationResponse threeoutput position
How does the same filter transformation move across patches while producing one output value per position?

What do you think happens?

A learned pattern appears in a different local region of an input. Does the convolution layer need an entirely new learned pattern for that new location?

  • Yes, because each location is a separate problem
  • No, because the same local transformation is reused
Reveal answer

Answer: No, because the same local transformation is reused.

Reusing the transformation lets the network search for the learned pattern at multiple locations.

Location and Pattern Hierarchies

Convolution layers learn translation-invariant patterns. If a network learns a pattern in one part of a picture, it can recognize that pattern elsewhere because the same local transformation is applied at different locations. The pattern does not have to be learned as an entirely new pattern for every position.

same transformationsame transformationInput patternone locationResponse mapone response locationInput patternanother locationResponse mapanother response location
What changes in the response map when a learned pattern shifts to another location in the input?

Convolutional networks also learn spatial hierarchies of patterns. An early convolution layer can learn small local patterns such as edges. A later layer can learn larger patterns made from features learned earlier. Further layers can continue combining these features into increasingly complex and abstract visual concepts. The important progression is from small and local to larger and more abstract.

provide building blockscombine intoLocal patternsedgesLarger patternscombined earlier featuresAbstract conceptsincreasing complexity
How do local responses become increasingly larger and more abstract patterns across layers?

Common Misunderstandings

  • Thinking convolution examines the entire image as one undivided pattern.

    Convolution examines small local regions or patches of the input.

    Fix: Track the operation from one local patch to another and focus on how responses are recorded across positions.

  • Assuming a new transformation is learned for every spatial location.

    The same learned transformation is reused across patches.

    Fix: Separate the changing patch location from the reused filter transformation.

  • Treating depth as another spatial direction.

    Height and width describe spatial position; depth represents channels.

    Fix: Interpret depth as the channel axis.

  • Assuming output channels are still specifically RGB color channels.

    Output channels represent filters and the responses produced by those filters.

    Fix: Read an output channel as a response map for one filter.

Check Your Understanding

MEDIUM

Describe the path from an input feature map to one output response map. Your explanation should mention a local patch, a filter, reuse across locations, spatial position, and output channels.

Hints
  • Start with what the convolution layer selects from the input.
  • Explain what remains the same as the selected patch changes.
  • Finish by describing what one output channel represents.
EASY

A pattern moves from one location in an input picture to another. Explain why a convolutional network can seek the same learned pattern at the new location without learning it as an entirely new pattern.

Hints
  • Focus on whether the transformation is tied to one location.
  • Use the term translation-invariant pattern in your explanation.

Key Takeaways

  1. Convolution examines small local patches instead of treating the entire image as one undivided pattern.
  2. The same learned transformation is reused across different patches, producing responses at different spatial positions.
  3. Height and width describe spatial position, while depth represents channels.
  4. Each filter can produce an output channel that acts as a two-dimensional response map.
  5. Convolutional networks recognize patterns across locations and build increasingly complex concepts from small local features.

Key Takeaways

  • A convolution operation extracts and examines local patches of an input feature map.
  • One learned transformation is reused across patches, allowing the same pattern to be recognized at multiple locations.
  • Feature maps have height, width, and depth axes; depth represents channels.
  • A filter produces a response map, which becomes a channel in the output feature map.
  • Convolutional networks progress from small local patterns to larger and more abstract spatial concepts.