Building a Convolutional Neural Network
Max pooling replaces each local window with its maximum value.
Try it: Pooling
How max or average pooling slides a window over a feature map and keeps one number per window, shrinking the map.
How it works
- Place the window over the top-left corner of the feature map.
- Max pooling keeps the largest value in the window; average pooling keeps the mean.
- Move the window by the stride and repeat.
- Output size per side = ⌊(n − k) / s⌋ + 1; each channel is pooled separately.
Default run (11 steps): 6×6 feature map, 2×2 max pooling, stride 2. Output is ⌊(6 − 2) / 2⌋ + 1 = 3 per side. … Done: the 6×6 map shrank to 3×3 — each output keeps one strongest value per window.
Simplified: One small single-channel feature map of whole numbers.
Loading the simulation…
A Small Region, One Survivor
A convolutional neural network, also called a convnet, is a type of deep-learning model used in computer vision. One problem a convnet can address is image classification: assigning an image to a category. In a small-training-dataset setting, the network must extract useful visual evidence while controlling how much information later layers need to process. Max pooling helps with that control by replacing each small local region with its maximum value.
How a Pooling Window Summarizes Evidence
Max pooling examines a small region of a feature map and carries forward the largest value in that region. The output therefore contains one value where the input contained several nearby values. In the source's interpretation, feature maps can represent the spatial presence of a pattern across different tiles. Keeping the maximum emphasizes the strongest local presence of that pattern.
Pooling one local window
A 2 × 2 local region contains the values 1, 7, 3, and 5. What value does max pooling produce for this region?
Inspect the region: The pooling operation considers only the four values inside this local 2 × 2 window.
Select the maximum: Among 1, 7, 3, and 5, the largest value is 7.
Write one output value: The four input values are represented by the single output value 7.
The pooled output for this window is 7.
Why Stride 2 Shrinks the Map
The window size determines how large a local area is examined. The stride determines how far the window advances before producing the next output. With the common 2 × 2 window and stride 2 arrangement, the operation produces fewer window positions along both height and width. Each spatial dimension is reduced by a factor of 2. For example, a 26 × 26 feature map becomes 13 × 13 after max pooling.
Following the spatial dimensions
A feature map has spatial dimensions of 26 × 26. It passes through max pooling with a 2 × 2 window and stride 2.
Consider the height: The stride-2 arrangement reduces the height by a factor of 2, changing 26 rows to 13 rows.
Consider the width: The same arrangement reduces the width by a factor of 2, changing 26 columns to 13 columns.
Combine the dimensions: The resulting spatial dimensions are 13 × 13.
The 26 × 26 feature map becomes a 13 × 13 feature map.
Channels Stay Separate
Pooling changes the spatial size of a feature map, but it does not remove the channel dimension. Max pooling is applied independently within each channel. A local region in one channel is summarized using that channel's own values; values from another channel are not included in the same maximum. The result keeps one downsampled spatial map for every original channel.
Why Downsampling Matters Later
Downsampling has two linked benefits. First, it reduces the number of feature-map coefficients that later layers must process. The source gives a 22 × 22 × 64 feature map containing 30,976 coefficients per sample. Flattening that map before a Dense layer of size 512 would create 15.8 million parameters, which would be far too large for the small model being discussed and could cause intense overfitting.
Second, repeated downsampling supports a spatial hierarchy. Later convolutional windows can cover an increasingly larger fraction of the original input. This matters for tasks such as recognizing a whole digit, because later features need access to information from a larger portion of the input rather than only a small local region.
For image classification, the network may need to recognize a whole digit rather than preserve every original spatial position forever. Downsampling lets later features use information from a broader portion of the input while reducing the amount of feature-map data that later layers process.
Three Ways to Reduce Spatial Size
Max pooling is one way to reduce spatial dimensions. A convolutional layer can also perform reduction by using strides, and a network can use average pooling. These operations should not be treated as interchangeable descriptions. Max pooling keeps the maximum value from each local region. Average pooling keeps the average value of each channel over that region. A strided convolution uses a convolutional layer before the reduction, with strides controlling how the convolution advances.
| Operation | What it carries forward or uses | Key distinction |
|---|---|---|
| Max pooling | The maximum value in each local region | Emphasizes the strongest local presence |
| Average pooling | The average value of each channel over the local region | Can weaken or dilute a strong local presence |
| Strided convolution | A convolutional layer using strides | Reduces spatial dimensions through a convolutional operation |
Common Pooling Mistakes
Treating max pooling as a search for one maximum across the entire feature map.
Max pooling makes separate local decisions for each window.
Fix:
Apply the maximum operation independently to every local window.Assuming pooling combines channels.
Pooling operates independently by channel.
Fix:
Pool the spatial regions within each channel separately.Explaining the size reduction without mentioning stride.
The window defines the local area, while the stride controls how far the window advances.
Fix:
For the common 2 × 2, stride 2 setup, connect the factor-of-two reduction to the stride-2 movement.Calling average pooling and max pooling the same operation.
Average pooling replaces the region with an average, while max pooling keeps its maximum.
Fix:
Name the aggregation used by the operation.Focusing only on shrinking the map and ignoring the spatial hierarchy.
Repeated downsampling also helps later windows cover a broader fraction of the original input.
Fix:
Explain both reduced processing cost and broader spatial features.
Convnet Use Cases
Convnets are used in computer vision applications. Image classification is one such problem: the input is an image, and the task is to assign that image to a category. Image-classification problems with small training datasets are identified as the most common use case when the organization is not a large technology company.
Framing a visual task
A computer receives an image and must assign it to a category using a relatively small training dataset. How should this task be described?
Identify the data: The input is visual data: an image.
Identify the task: The computer must assign the image to a category, so the task is image classification.
Identify the model family: A convolutional neural network, or convnet, is a deep-learning model used for this kind of computer vision work.
Identify the practical setting: The training dataset is small, which is an important use case described in the source.
This is a computer-vision image-classification problem using a convnet in a small-training-dataset setting.
Practice and Recall
A feature map has spatial dimensions of 26 × 26 and is processed with a 2 × 2 max-pooling window and stride 2. Explain the new spatial dimensions, describe what happens to the channels, and state what value each local window contributes to the output.
Hints
- Apply the factor-of-two reduction separately to height and width.
- Pooling changes spatial size but does not remove the channel dimension.
- Each local window contributes its maximum value.
A team wants to reduce spatial dimensions but is deciding between max pooling, average pooling, and a strided convolution. Explain the defining operation for each choice and why max pooling may preserve strong evidence of a local pattern.
Hints
- Compare maximum, average, and convolutional processing.
- Mention the role of stride for the convolutional option.
- Relate the maximum to the strongest local presence of a pattern.
Key Takeaways
- Max pooling replaces each local window with its maximum value, preserving the strongest local activation.
- A 2 × 2 window with stride 2 reduces both spatial dimensions by a factor of 2, such as 26 × 26 becoming 13 × 13.
- Pooling is performed independently for each channel, so spatial size changes while the channel dimension remains.
- Downsampling reduces later processing cost and supports a hierarchy of features covering broader regions of the original input.
- A convnet is a deep-learning model used in computer vision, including image classification with small training datasets.