Concepts / Spatial Hierarchies in Computer Vision

Spatial Hierarchies in Computer Vision

Max pooling replaces each local window with its maximum value.

  • Programming

The Strongest Local Evidence

A convolutional neural network does not need to preserve every spatial position forever. After earlier layers have detected useful patterns, the network can compress nearby positions while keeping the strongest evidence that a pattern is present. Max pooling performs this compression: it examines small regions of a feature map and carries forward the largest value from each region.

One Window, One Output

Suppose a max-pooling window covers four values in a 2 × 2 region. The operation compares those four values and writes only the largest one into the output feature map. The other values in that local region do not become separate output values. In this way, a group of nearby spatial positions is represented by one value: the strongest local activation for that channel.

Reducing four local values to one

A 2 × 2 max-pooling window contains the values 2, 7, 4, and 5. What value does this window contribute to the pooled feature map?

Inspect the local region: The window contains four values: 2, 7, 4, and 5.

Select the maximum: Among those four values, 7 is the largest.

Write one output: The entire 2 × 2 region is represented by the single output value 7.

The window produces the pooled value 7.

maximummaximumWindow A2, 7, 4, 57pooled valueWindow B1, 3, 6, 26pooled value
How do the values inside separate 2 × 2 windows map to single maximum values in the pooled feature map?

Stride and Spatial Size

The window size tells max pooling how large a local region to inspect. The stride tells it how far to move before producing the next output. A common configuration uses a 2 × 2 window with stride 2. The window therefore visits non-overlapping 2 × 2 regions, moving two positions at a time horizontally and vertically.

Why both dimensions are halved

A feature map has spatial dimensions 26 × 26. Max pooling uses a 2 × 2 window with stride 2. What are the resulting spatial dimensions?

Move across the width: With stride 2, the window advances through the width in steps of two, so the number of horizontal positions is reduced by a factor of 2.

Move down the height: The same movement happens vertically, so the number of vertical positions is also reduced by a factor of 2.

State the output size: The 26 × 26 spatial area becomes 13 × 13.

The pooled feature map has spatial dimensions 13 × 13.

visit local regionsone maximum per windowInput feature map26 × 262 × 2 windowsstride 2Pooled feature map13 × 13
Which non-overlapping input windows are visited, and how do their positions determine the smaller output height and width?

Channels Stay Separate

A feature map can contain multiple channels. Max pooling is applied independently within each channel. The values in one channel are compared only with other values from that same channel; pooling does not search across the entire feature map and does not select a value from one channel to represent all channels.

same channelsame channellocal maximalocal maximaChannel 1local valuesMax poolingChannel 1Channel 1smaller spatial mapChannel 2local valuesMax poolingChannel 2Channel 2smaller spatial map
How does the same pooling operation transform each channel separately without mixing values between channels?

When reasoning about a pooled tensor, track two questions separately: how many spatial positions remain, and how many channels remain. In the described pooling operation, the spatial dimensions shrink while the channel dimension remains present.

From Smaller Maps to Broader Features

Downsampling has two linked benefits. It reduces the number of feature-map coefficients that later layers must process, and it supports a hierarchy of increasingly broad spatial features. After repeated downsampling, a later convolutional window covers a larger fraction of the original input than a similarly sized window would cover if the feature map stayed large.

Why coefficient count matters

Consider a 22 × 22 × 64 feature map. The source describes this map as containing 30,976 coefficients per sample. What problem can arise if it is flattened before a Dense layer of size 512?

Count the incoming coefficients: The feature map contains 22 × 22 × 64, or 30,976, coefficients per sample.

Connect to the Dense layer: Flattening sends those coefficients toward a Dense layer with 512 units.

Consider the resulting parameter count: The source reports that this connection would create 15.8 million parameters.

Interpret the consequence: That parameter count would be far too large for the small model being discussed and could cause intense overfitting.

Downsampling helps control the number of coefficients that later layers must process.

2 × 2 max pooling, stride 2fewer spatial positions to processFeature map26 × 26Feature map13 × 13Broader spatialfeatureslater layers
How does a feature map become smaller through pooling, and why does processing fewer spatial positions help later computation?

The hierarchy is useful for tasks such as recognizing a whole digit. If every feature map stayed large, a 3 × 3 window in a later layer would still describe only a relatively small region of the original input. Downsampling allows later features to draw on information from across a larger portion of that input.

Pooling Alternatives

Max pooling is one way to reduce spatial dimensions, but it is not the only one. Average pooling replaces each local region with the average value of that channel over the region rather than the maximum. A convolutional layer can also reduce spatial dimensions by using strides. These operations should not be treated as interchangeable: they use different mechanisms to summarize or transform local information.

MethodLocal resultHow spatial reduction occursChannel handling
Max poolingThe maximum value in each local regionA window and stride determine the visited regions; 2 × 2 with stride 2 halves each spatial dimensionApplied independently by channel
Average poolingThe average value in each local regionA pooling window and its movement determine the reduced spatial mapThe average is taken for each channel over its local region
Strided convolutionA convolutional transformationA convolutional layer uses strides to reduce spatial dimensionsThe source pack identifies it as another dimensionality-reduction option but does not describe channel behavior here

Mistakes in Dimension Reasoning

  • Taking the maximum over the entire feature map

    Max pooling makes many local decisions. It replaces each local window with its maximum rather than producing one value for the whole feature map.

    Fix: Identify the window boundaries first, then select one maximum for each window.

  • Mixing values from different channels

    Pooling operates independently by channel.

    Fix: Pool within each channel and preserve the channel dimension.

  • Assuming a 2 × 2 window always means a two-position overlap

    The stride controls how far the window advances. With stride 2, the common setup visits non-overlapping regions.

    Fix: Track the window movement separately from the window size.

  • Thinking downsampling only changes notation

    Downsampling reduces the number of feature-map coefficients that later layers must process.

    Fix: Recalculate the spatial dimensions after each reduction and connect them to the later computational workload.

  • Confusing max pooling with average pooling

    Max pooling preserves the largest local value, whereas average pooling uses the average value for the channel over that region.

    Fix: Ask which local summary the operation is specified to produce.

Check Your Understanding

EASY

A single channel contains a 4 × 4 feature map. Apply max pooling with a 2 × 2 window and stride 2. The four local regions contain, in reading order, the values 1, 8, 3, 2; 5, 4, 6, 1; 7, 0, 2, 3; and 4, 9, 1, 5. What four values appear in the 2 × 2 pooled map, and why does the output have fewer spatial positions?

Hints
  • Treat each group of four values as one local window.
  • Select the largest value from each group.
  • A 2 × 2 window with stride 2 reduces each spatial dimension by a factor of 2.

What do you think happens?

A feature map has two channels. If one channel contains a local maximum of 9 and the other contains a local maximum of 4 in corresponding windows, will max pooling output only 9 for both channels?

  • Yes, because pooling chooses one value for the whole feature map
  • No, each channel keeps its own local maximum
  • Yes, because the larger value replaces the smaller channel value
  • No, pooling averages the two channel maxima
Reveal answer

Answer: No, each channel keeps its own local maximum.

Max pooling operates independently by channel. It does not compare values across channels.

What to Remember

  1. Max pooling replaces each local window with its maximum value.
  2. A 2 × 2 window with stride 2 visits non-overlapping regions and reduces height and width by a factor of 2.
  3. Pooling is performed independently for each channel, so spatial reduction does not mean selecting one value across all channels.
  4. Downsampling reduces the number of coefficients later layers process and helps build increasingly broad spatial features.
  5. Average pooling uses local averages, while strided convolution uses strides in a convolutional layer to reduce spatial dimensions.

Key Takeaways

  • Max pooling converts each local region into one output by preserving the region's largest value.
  • With a 2 × 2 window and stride 2, each spatial dimension is reduced by a factor of 2.
  • The operation is performed separately in every channel rather than across the entire feature map.
  • Downsampling lowers the number of coefficients and helps later layers represent broader regions of the original input.
  • Max pooling, average pooling, and strided convolution reduce spatial information in different ways.