Feature Maps in Convolutional Neural Networks
Max pooling replaces each local window with its maximum value.
Try it: Pooling
How max or average pooling slides a window over a feature map and keeps one number per window, shrinking the map.
How it works
- Place the window over the top-left corner of the feature map.
- Max pooling keeps the largest value in the window; average pooling keeps the mean.
- Move the window by the stride and repeat.
- Output size per side = ⌊(n − k) / s⌋ + 1; each channel is pooled separately.
Default run (11 steps): 6×6 feature map, 2×2 max pooling, stride 2. Output is ⌊(6 − 2) / 2⌋ + 1 = 3 per side. … Done: the 6×6 map shrank to 3×3 — each output keeps one strongest value per window.
Simplified: One small single-channel feature map of whole numbers.
Loading the simulation…
The Strongest Local Evidence
A convolutional neural network does not need to preserve every spatial position forever. After earlier layers have detected useful patterns, the network can compress nearby positions while keeping the strongest evidence that a pattern is present. Max pooling performs this compression by examining small regions of a feature map and carrying forward the largest value from each region.
Max pooling makes many local decisions. It does not search for one maximum across the entire feature map, and it does not compare values across channels.
One Window One Output
A max-pooling window covers a local region of one channel of a feature map. The operation compares the values inside that region and writes the largest one into a single output position. The output therefore records the strongest local activation found by that window.
Selecting a local maximum
Apply max pooling to this 2 × 2 region: 1, 4, 2, 3.
Inspect the region: The pooling window sees the four values 1, 4, 2, and 3.
Compare the values: The largest value in this local region is 4.
Write one output: The entire 2 × 2 region is represented by the single output value 4.
The local region produces one pooled value: 4.
Window Movement and Size Reduction
The window size determines how much local area is examined at once. The stride determines how far the window advances before producing its next output. In the common 2 × 2, stride 2 setup, the window advances two positions at a time horizontally and vertically. This arrangement produces one output for each non-overlapping 2 × 2 region, reducing both spatial dimensions by a factor of 2.
Reducing a feature map
A feature map has spatial dimensions 26 × 26. What are its spatial dimensions after 2 × 2 max pooling with stride 2?
Apply the height reduction: The common 2 × 2, stride 2 arrangement reduces the height by a factor of 2: 26 becomes 13.
Apply the width reduction: The same arrangement reduces the width by a factor of 2: 26 becomes 13.
Keep the channel dimension: Pooling changes the spatial size but does not remove the channel dimension.
The 26 × 26 spatial map becomes 13 × 13 for each channel.
Independent Channel Processing
A feature map can contain multiple channels, with each channel representing a different learned response. Max pooling processes each channel independently. The same local pooling rule is applied within one channel at a time, so values from different channels are not compared with one another. The spatial dimensions become smaller, but the channel dimension remains.
Choosing one maximum across the entire feature map
Max pooling makes many local decisions within each channel. It does not select one global maximum.
Fix:
Apply the window separately at each spatial location within each channel.Comparing values across channels
Pooling operates independently by channel rather than selecting one value across the whole feature map.
Fix:
Preserve a separate pooled output map for every channel.
Cost and Spatial Hierarchy
Downsampling has two linked benefits. First, it reduces the number of feature-map coefficients that later layers must process. The source gives a 22 × 22 × 64 feature map containing 30,976 coefficients per sample. Flattening that map before a Dense layer of size 512 would create 15.8 million parameters, which would be far too large for the small model being discussed and could cause intense overfitting.
Second, repeated downsampling supports a spatial hierarchy. Later convolutional windows can cover an increasingly larger fraction of the original input. If feature maps remained large throughout the network, a 3 × 3 window in a later layer would still describe only a relatively small region of the original input. For tasks such as recognizing a whole digit, later features need access to information from across a larger portion of the input.
Three Ways to Reduce Space
| Operation | What it does to each local region | How the source distinguishes it |
|---|---|---|
| Max pooling | Keeps the maximum value | Emphasizes the strongest local presence of a feature |
| Average pooling | Replaces the region with its average value for each channel | May weaken or dilute strong local evidence |
| Strided convolution | Uses a convolutional layer before or during spatial reduction through strides | Reduces spatial dimensions through a convolutional layer rather than pooling |
Max pooling is useful when the strongest local activation is the important signal. The source describes feature maps as encoding the spatial presence of a pattern or concept across different tiles. Keeping the maximum emphasizes the strongest local presence. Average pooling instead combines the values into an average, which can dilute that evidence. A convolutional layer can also reduce spatial dimensions by using strides, but that reduction occurs through a convolutional layer rather than through a pooling summary.
Common Interpretation Errors
Thinking max pooling preserves every value in a region
The window is replaced by one value: the maximum.
Fix:
Track one output value for each local window.Assuming stride 2 means only the values are reduced
The stride controls how far the window advances, and the common 2 × 2, stride 2 setup reduces each spatial dimension by a factor of 2.
Fix:
Track both the window size and its movement across height and width.Assuming pooling removes channels
Pooling changes spatial size but operates independently for each channel.
Fix:
Keep one smaller spatial map for each original channel.Treating max pooling and average pooling as equivalent
Max pooling keeps the strongest value, while average pooling can weaken or dilute strong evidence.
Fix:
Identify whether the operation selects a maximum or calculates an average.
Check Your Reasoning
A feature map has spatial dimensions 12 × 12 and several channels. It passes through 2 × 2 max pooling with stride 2. Explain what happens to the height, width, and channels. Then explain what value a single pooling window contributes to one channel.
Hints
- Apply the factor-of-two reduction separately to height and width.
- The channel dimension is preserved.
- For one channel, compare the values inside one local window and retain the largest.
A later CNN layer must recognize a whole digit. Explain why repeated downsampling can help that layer access information from a larger portion of the original input while also reducing the number of coefficients processed by later layers.
Hints
- Connect smaller feature maps with fewer coefficients.
- Consider what fraction of the original input a later convolutional window can describe.
- Use the idea of a spatial hierarchy.
Key Takeaways
- Max pooling replaces each local window with its maximum value, preserving the strongest local activation.
- A 2 × 2 window with stride 2 reduces height and width by a factor of 2.
- Pooling operates independently within each channel, so it does not compare values across channels.
- Downsampling reduces the number of coefficients later layers process and supports increasingly broad spatial features.
- Average pooling summarizes a region with an average, while strided convolution reduces spatial dimensions through a convolutional layer.
Key Takeaways
- Max pooling converts each local region into one output containing the region's maximum value.
- With a 2 × 2 window and stride 2, each spatial dimension is reduced by a factor of 2.
- The operation is repeated independently for every channel.
- Downsampling lowers later computational demands and helps later layers represent broader spatial features.
- Max pooling, average pooling, and strided convolution reduce spatial dimensions through different local operations.