Spatial Hierarchies in Computer Vision
Max pooling replaces each local window with its maximum value.
The Strongest Local Evidence
A convolutional neural network does not need to preserve every spatial position forever. After earlier layers have detected useful patterns, the network can compress nearby positions while keeping the strongest evidence that a pattern is present. Max pooling performs this compression: it examines small regions of a feature map and carries forward the largest value from each region.
One Window, One Output
Suppose a max-pooling window covers four values in a 2 × 2 region. The operation compares those four values and writes only the largest one into the output feature map. The other values in that local region do not become separate output values. In this way, a group of nearby spatial positions is represented by one value: the strongest local activation for that channel.
Reducing four local values to one
A 2 × 2 max-pooling window contains the values 2, 7, 4, and 5. What value does this window contribute to the pooled feature map?
Inspect the local region: The window contains four values: 2, 7, 4, and 5.
Select the maximum: Among those four values, 7 is the largest.
Write one output: The entire 2 × 2 region is represented by the single output value 7.
The window produces the pooled value 7.
Stride and Spatial Size
The window size tells max pooling how large a local region to inspect. The stride tells it how far to move before producing the next output. A common configuration uses a 2 × 2 window with stride 2. The window therefore visits non-overlapping 2 × 2 regions, moving two positions at a time horizontally and vertically.
Why both dimensions are halved
A feature map has spatial dimensions 26 × 26. Max pooling uses a 2 × 2 window with stride 2. What are the resulting spatial dimensions?
Move across the width: With stride 2, the window advances through the width in steps of two, so the number of horizontal positions is reduced by a factor of 2.
Move down the height: The same movement happens vertically, so the number of vertical positions is also reduced by a factor of 2.
State the output size: The 26 × 26 spatial area becomes 13 × 13.
The pooled feature map has spatial dimensions 13 × 13.
Channels Stay Separate
A feature map can contain multiple channels. Max pooling is applied independently within each channel. The values in one channel are compared only with other values from that same channel; pooling does not search across the entire feature map and does not select a value from one channel to represent all channels.
When reasoning about a pooled tensor, track two questions separately: how many spatial positions remain, and how many channels remain. In the described pooling operation, the spatial dimensions shrink while the channel dimension remains present.
From Smaller Maps to Broader Features
Downsampling has two linked benefits. It reduces the number of feature-map coefficients that later layers must process, and it supports a hierarchy of increasingly broad spatial features. After repeated downsampling, a later convolutional window covers a larger fraction of the original input than a similarly sized window would cover if the feature map stayed large.
Why coefficient count matters
Consider a 22 × 22 × 64 feature map. The source describes this map as containing 30,976 coefficients per sample. What problem can arise if it is flattened before a Dense layer of size 512?
Count the incoming coefficients: The feature map contains 22 × 22 × 64, or 30,976, coefficients per sample.
Connect to the Dense layer: Flattening sends those coefficients toward a Dense layer with 512 units.
Consider the resulting parameter count: The source reports that this connection would create 15.8 million parameters.
Interpret the consequence: That parameter count would be far too large for the small model being discussed and could cause intense overfitting.
Downsampling helps control the number of coefficients that later layers must process.
The hierarchy is useful for tasks such as recognizing a whole digit. If every feature map stayed large, a 3 × 3 window in a later layer would still describe only a relatively small region of the original input. Downsampling allows later features to draw on information from across a larger portion of that input.
Pooling Alternatives
Max pooling is one way to reduce spatial dimensions, but it is not the only one. Average pooling replaces each local region with the average value of that channel over the region rather than the maximum. A convolutional layer can also reduce spatial dimensions by using strides. These operations should not be treated as interchangeable: they use different mechanisms to summarize or transform local information.
| Method | Local result | How spatial reduction occurs | Channel handling |
|---|---|---|---|
| Max pooling | The maximum value in each local region | A window and stride determine the visited regions; 2 × 2 with stride 2 halves each spatial dimension | Applied independently by channel |
| Average pooling | The average value in each local region | A pooling window and its movement determine the reduced spatial map | The average is taken for each channel over its local region |
| Strided convolution | A convolutional transformation | A convolutional layer uses strides to reduce spatial dimensions | The source pack identifies it as another dimensionality-reduction option but does not describe channel behavior here |
Mistakes in Dimension Reasoning
Taking the maximum over the entire feature map
Max pooling makes many local decisions. It replaces each local window with its maximum rather than producing one value for the whole feature map.
Fix:
Identify the window boundaries first, then select one maximum for each window.Mixing values from different channels
Pooling operates independently by channel.
Fix:
Pool within each channel and preserve the channel dimension.Assuming a 2 × 2 window always means a two-position overlap
The stride controls how far the window advances. With stride 2, the common setup visits non-overlapping regions.
Fix:
Track the window movement separately from the window size.Thinking downsampling only changes notation
Downsampling reduces the number of feature-map coefficients that later layers must process.
Fix:
Recalculate the spatial dimensions after each reduction and connect them to the later computational workload.Confusing max pooling with average pooling
Max pooling preserves the largest local value, whereas average pooling uses the average value for the channel over that region.
Fix:
Ask which local summary the operation is specified to produce.
Check Your Understanding
A single channel contains a 4 × 4 feature map. Apply max pooling with a 2 × 2 window and stride 2. The four local regions contain, in reading order, the values 1, 8, 3, 2; 5, 4, 6, 1; 7, 0, 2, 3; and 4, 9, 1, 5. What four values appear in the 2 × 2 pooled map, and why does the output have fewer spatial positions?
Hints
- Treat each group of four values as one local window.
- Select the largest value from each group.
- A 2 × 2 window with stride 2 reduces each spatial dimension by a factor of 2.
What do you think happens?
A feature map has two channels. If one channel contains a local maximum of 9 and the other contains a local maximum of 4 in corresponding windows, will max pooling output only 9 for both channels?
Reveal answer
Answer: No, each channel keeps its own local maximum.
Max pooling operates independently by channel. It does not compare values across channels.
What to Remember
- Max pooling replaces each local window with its maximum value.
- A 2 × 2 window with stride 2 visits non-overlapping regions and reduces height and width by a factor of 2.
- Pooling is performed independently for each channel, so spatial reduction does not mean selecting one value across all channels.
- Downsampling reduces the number of coefficients later layers process and helps build increasingly broad spatial features.
- Average pooling uses local averages, while strided convolution uses strides in a convolutional layer to reduce spatial dimensions.
Key Takeaways
- Max pooling converts each local region into one output by preserving the region's largest value.
- With a 2 × 2 window and stride 2, each spatial dimension is reduced by a factor of 2.
- The operation is performed separately in every channel rather than across the entire feature map.
- Downsampling lowers the number of coefficients and helps later layers represent broader regions of the original input.
- Max pooling, average pooling, and strided convolution reduce spatial information in different ways.