Concepts / Feature Maps in Convolutional Neural Networks

Feature Maps in Convolutional Neural Networks

Max pooling replaces each local window with its maximum value.

  • Programming
Interactive lab

Try it: Pooling

How max or average pooling slides a window over a feature map and keeps one number per window, shrinking the map.

How it works

  1. Place the window over the top-left corner of the feature map.
  2. Max pooling keeps the largest value in the window; average pooling keeps the mean.
  3. Move the window by the stride and repeat.
  4. Output size per side = ⌊(n − k) / s⌋ + 1; each channel is pooled separately.

Default run (11 steps): 6×6 feature map, 2×2 max pooling, stride 2. Output is ⌊(6 − 2) / 2⌋ + 1 = 3 per side. … Done: the 6×6 map shrank to 3×3 — each output keeps one strongest value per window.

Simplified: One small single-channel feature map of whole numbers.

Educational simulation

Loading the simulation…

The Strongest Local Evidence

A convolutional neural network does not need to preserve every spatial position forever. After earlier layers have detected useful patterns, the network can compress nearby positions while keeping the strongest evidence that a pattern is present. Max pooling performs this compression by examining small regions of a feature map and carrying forward the largest value from each region.

Max pooling makes many local decisions. It does not search for one maximum across the entire feature map, and it does not compare values across channels.

One Window One Output

A max-pooling window covers a local region of one channel of a feature map. The operation compares the values inside that region and writes the largest one into a single output position. The output therefore records the strongest local activation found by that window.

Selecting a local maximum

Apply max pooling to this 2 × 2 region: 1, 4, 2, 3.

Inspect the region: The pooling window sees the four values 1, 4, 2, and 3.

Compare the values: The largest value in this local region is 4.

Write one output: The entire 2 × 2 region is represented by the single output value 4.

The local region produces one pooled value: 4.

take the largest value2 × 2 region1, 4, 2, 34local maximum
How does a 2 × 2 max-pooling window select one maximum value from a local feature-map region?

Window Movement and Size Reduction

The window size determines how much local area is examined at once. The stride determines how far the window advances before producing its next output. In the common 2 × 2, stride 2 setup, the window advances two positions at a time horizontally and vertically. This arrangement produces one output for each non-overlapping 2 × 2 region, reducing both spatial dimensions by a factor of 2.

Reducing a feature map

A feature map has spatial dimensions 26 × 26. What are its spatial dimensions after 2 × 2 max pooling with stride 2?

Apply the height reduction: The common 2 × 2, stride 2 arrangement reduces the height by a factor of 2: 26 becomes 13.

Apply the width reduction: The same arrangement reduces the width by a factor of 2: 26 becomes 13.

Keep the channel dimension: Pooling changes the spatial size but does not remove the channel dimension.

The 26 × 26 spatial map becomes 13 × 13 for each channel.

2 × 2 windows across height2 × 2 windows across widthcombine height and widthcombine height and width26 × 26 mapinput positions13 output rowswindow advances by 213 × 13 mapper channel13 output columnswindow advances by 2
How do the pooling window positions move across the input, and why do they produce an output with half the height and width?

Independent Channel Processing

A feature map can contain multiple channels, with each channel representing a different learned response. Max pooling processes each channel independently. The same local pooling rule is applied within one channel at a time, so values from different channels are not compared with one another. The spatial dimensions become smaller, but the channel dimension remains.

separate pathseparate pathretain channel identityretain channel identityChannel Alocal valuesMax poolingwithin Channel AChannel Asmaller spatial mapChannel Blocal valuesMax poolingwithin Channel BChannel Bsmaller spatial map
How does the same pooling operation transform each channel separately without combining information between channels?
  • Choosing one maximum across the entire feature map

    Max pooling makes many local decisions within each channel. It does not select one global maximum.

    Fix: Apply the window separately at each spatial location within each channel.

  • Comparing values across channels

    Pooling operates independently by channel rather than selecting one value across the whole feature map.

    Fix: Preserve a separate pooled output map for every channel.

Cost and Spatial Hierarchy

Downsampling has two linked benefits. First, it reduces the number of feature-map coefficients that later layers must process. The source gives a 22 × 22 × 64 feature map containing 30,976 coefficients per sample. Flattening that map before a Dense layer of size 512 would create 15.8 million parameters, which would be far too large for the small model being discussed and could cause intense overfitting.

Second, repeated downsampling supports a spatial hierarchy. Later convolutional windows can cover an increasingly larger fraction of the original input. If feature maps remained large throughout the network, a 3 × 3 window in a later layer would still describe only a relatively small region of the original input. For tasks such as recognizing a whole digit, later features need access to information from across a larger portion of the input.

compress nearby positionsreduce coefficientsenable later windows to cover more inputLarge feature mapmany coefficientsDownsamplingsmaller height and widthFewer coefficientsless later processingBroader spatialfeatureslarger fraction of theinput
How does reducing feature-map height and width decrease later computation while supporting increasingly broad spatial information?

Three Ways to Reduce Space

OperationWhat it does to each local regionHow the source distinguishes it
Max poolingKeeps the maximum valueEmphasizes the strongest local presence of a feature
Average poolingReplaces the region with its average value for each channelMay weaken or dilute strong local evidence
Strided convolutionUses a convolutional layer before or during spatial reduction through stridesReduces spatial dimensions through a convolutional layer rather than pooling

Max pooling is useful when the strongest local activation is the important signal. The source describes feature maps as encoding the spatial presence of a pattern or concept across different tiles. Keeping the maximum emphasizes the strongest local presence. Average pooling instead combines the values into an average, which can dilute that evidence. A convolutional layer can also reduce spatial dimensions by using strides, but that reduction occurs through a convolutional layer rather than through a pooling summary.

select maximumcalculate averageapply convolution with stridespreserve strongest local presencesummarize by averageproduce reduced convolution outputLocal regionsame input areaMax poolinglargest valueStrongest evidencelocal maximumAverage poolingaverage valueAveraged evidencelocal averageStrided convolutionconvolution with stridesConvolution outputspatially reduced
What is the difference between selecting the maximum, averaging values, and using a strided convolution when reducing spatial dimensions?

Common Interpretation Errors

  • Thinking max pooling preserves every value in a region

    The window is replaced by one value: the maximum.

    Fix: Track one output value for each local window.

  • Assuming stride 2 means only the values are reduced

    The stride controls how far the window advances, and the common 2 × 2, stride 2 setup reduces each spatial dimension by a factor of 2.

    Fix: Track both the window size and its movement across height and width.

  • Assuming pooling removes channels

    Pooling changes spatial size but operates independently for each channel.

    Fix: Keep one smaller spatial map for each original channel.

  • Treating max pooling and average pooling as equivalent

    Max pooling keeps the strongest value, while average pooling can weaken or dilute strong evidence.

    Fix: Identify whether the operation selects a maximum or calculates an average.

Check Your Reasoning

EASY

A feature map has spatial dimensions 12 × 12 and several channels. It passes through 2 × 2 max pooling with stride 2. Explain what happens to the height, width, and channels. Then explain what value a single pooling window contributes to one channel.

Hints
  • Apply the factor-of-two reduction separately to height and width.
  • The channel dimension is preserved.
  • For one channel, compare the values inside one local window and retain the largest.
MEDIUM

A later CNN layer must recognize a whole digit. Explain why repeated downsampling can help that layer access information from a larger portion of the original input while also reducing the number of coefficients processed by later layers.

Hints
  • Connect smaller feature maps with fewer coefficients.
  • Consider what fraction of the original input a later convolutional window can describe.
  • Use the idea of a spatial hierarchy.

Key Takeaways

  1. Max pooling replaces each local window with its maximum value, preserving the strongest local activation.
  2. A 2 × 2 window with stride 2 reduces height and width by a factor of 2.
  3. Pooling operates independently within each channel, so it does not compare values across channels.
  4. Downsampling reduces the number of coefficients later layers process and supports increasingly broad spatial features.
  5. Average pooling summarizes a region with an average, while strided convolution reduces spatial dimensions through a convolutional layer.

Key Takeaways

  • Max pooling converts each local region into one output containing the region's maximum value.
  • With a 2 × 2 window and stride 2, each spatial dimension is reduced by a factor of 2.
  • The operation is repeated independently for every channel.
  • Downsampling lowers later computational demands and helps later layers represent broader spatial features.
  • Max pooling, average pooling, and strided convolution reduce spatial dimensions through different local operations.