Concepts / Building a Convolutional Neural Network

Building a Convolutional Neural Network

Max pooling replaces each local window with its maximum value.

  • Programming
Interactive lab

Try it: Pooling

How max or average pooling slides a window over a feature map and keeps one number per window, shrinking the map.

How it works

  1. Place the window over the top-left corner of the feature map.
  2. Max pooling keeps the largest value in the window; average pooling keeps the mean.
  3. Move the window by the stride and repeat.
  4. Output size per side = ⌊(n − k) / s⌋ + 1; each channel is pooled separately.

Default run (11 steps): 6×6 feature map, 2×2 max pooling, stride 2. Output is ⌊(6 − 2) / 2⌋ + 1 = 3 per side. … Done: the 6×6 map shrank to 3×3 — each output keeps one strongest value per window.

Simplified: One small single-channel feature map of whole numbers.

Educational simulation

Loading the simulation…

A Small Region, One Survivor

A convolutional neural network, also called a convnet, is a type of deep-learning model used in computer vision. One problem a convnet can address is image classification: assigning an image to a category. In a small-training-dataset setting, the network must extract useful visual evidence while controlling how much information later layers need to process. Max pooling helps with that control by replacing each small local region with its maximum value.

How a Pooling Window Summarizes Evidence

Max pooling examines a small region of a feature map and carries forward the largest value in that region. The output therefore contains one value where the input contained several nearby values. In the source's interpretation, feature maps can represent the spatial presence of a pattern across different tiles. Keeping the maximum emphasizes the strongest local presence of that pattern.

take maximum2 × 2 region1, 7, 3, 57local maximum
How does a 2 × 2 region of feature-map values become one output value?

Pooling one local window

A 2 × 2 local region contains the values 1, 7, 3, and 5. What value does max pooling produce for this region?

Inspect the region: The pooling operation considers only the four values inside this local 2 × 2 window.

Select the maximum: Among 1, 7, 3, and 5, the largest value is 7.

Write one output value: The four input values are represented by the single output value 7.

The pooled output for this window is 7.

Why Stride 2 Shrinks the Map

The window size determines how large a local area is examined. The stride determines how far the window advances before producing the next output. With the common 2 × 2 window and stride 2 arrangement, the operation produces fewer window positions along both height and width. Each spatial dimension is reduced by a factor of 2. For example, a 26 × 26 feature map becomes 13 × 13 after max pooling.

poolpoolpoolpoolInput rows 1–2Input columns 1–2Output row 1Output column 1Input rows 1–2Input columns 3–4Output row 1Output column 2Input rows 3–4Input columns 1–2Output row 2Output column 1Input rows 3–4Input columns 3–4Output row 2Output column 2
How do the window positions map from the input feature map to fewer output rows and columns?

Following the spatial dimensions

A feature map has spatial dimensions of 26 × 26. It passes through max pooling with a 2 × 2 window and stride 2.

Consider the height: The stride-2 arrangement reduces the height by a factor of 2, changing 26 rows to 13 rows.

Consider the width: The same arrangement reduces the width by a factor of 2, changing 26 columns to 13 columns.

Combine the dimensions: The resulting spatial dimensions are 13 × 13.

The 26 × 26 feature map becomes a 13 × 13 feature map.

Channels Stay Separate

Pooling changes the spatial size of a feature map, but it does not remove the channel dimension. Max pooling is applied independently within each channel. A local region in one channel is summarized using that channel's own values; values from another channel are not included in the same maximum. The result keeps one downsampled spatial map for every original channel.

own valuesown valuesown valuesreduce spatial sizereduce spatial sizereduce spatial sizeChannel Alocal windowsMax pool Aspatial reductionChannel A outputdownsampled mapChannel Blocal windowsMax pool Bspatial reductionChannel B outputdownsampled mapChannel Clocal windowsMax pool Cspatial reductionChannel C outputdownsampled map
How does the same pooling operation transform each channel without combining values across channels?

Why Downsampling Matters Later

Downsampling has two linked benefits. First, it reduces the number of feature-map coefficients that later layers must process. The source gives a 22 × 22 × 64 feature map containing 30,976 coefficients per sample. Flattening that map before a Dense layer of size 512 would create 15.8 million parameters, which would be far too large for the small model being discussed and could cause intense overfitting.

Second, repeated downsampling supports a spatial hierarchy. Later convolutional windows can cover an increasingly larger fraction of the original input. This matters for tasks such as recognizing a whole digit, because later features need access to information from a larger portion of the input rather than only a small local region.

compress nearby positionsreduce spatial sizesupport later layersLarge feature mapmany coefficientsDownsamplingfewer spatial positionsSmaller feature mapfewer coefficientsBroader spatialfeatureslarger input regions
How does reducing spatial dimensions change the amount of data processed and allow later layers to represent larger regions?

For image classification, the network may need to recognize a whole digit rather than preserve every original spatial position forever. Downsampling lets later features use information from a broader portion of the input while reducing the amount of feature-map data that later layers process.

Three Ways to Reduce Spatial Size

Max pooling is one way to reduce spatial dimensions. A convolutional layer can also perform reduction by using strides, and a network can use average pooling. These operations should not be treated as interchangeable descriptions. Max pooling keeps the maximum value from each local region. Average pooling keeps the average value of each channel over that region. A strided convolution uses a convolutional layer before the reduction, with strides controlling how the convolution advances.

select maximumcalculate averageadvance convolution by strideLocal regionmaximumMax poolingstrongest valueLocal regionaverageAverage poolingaverage valueLocal inputconvolutional layerStrided convolutionstride-based reduction
What changes when a local region is summarized by a maximum, an average, or a strided convolution?
OperationWhat it carries forward or usesKey distinction
Max poolingThe maximum value in each local regionEmphasizes the strongest local presence
Average poolingThe average value of each channel over the local regionCan weaken or dilute a strong local presence
Strided convolutionA convolutional layer using stridesReduces spatial dimensions through a convolutional operation

Common Pooling Mistakes

  • Treating max pooling as a search for one maximum across the entire feature map.

    Max pooling makes separate local decisions for each window.

    Fix: Apply the maximum operation independently to every local window.

  • Assuming pooling combines channels.

    Pooling operates independently by channel.

    Fix: Pool the spatial regions within each channel separately.

  • Explaining the size reduction without mentioning stride.

    The window defines the local area, while the stride controls how far the window advances.

    Fix: For the common 2 × 2, stride 2 setup, connect the factor-of-two reduction to the stride-2 movement.

  • Calling average pooling and max pooling the same operation.

    Average pooling replaces the region with an average, while max pooling keeps its maximum.

    Fix: Name the aggregation used by the operation.

  • Focusing only on shrinking the map and ignoring the spatial hierarchy.

    Repeated downsampling also helps later windows cover a broader fraction of the original input.

    Fix: Explain both reduced processing cost and broader spatial features.

Convnet Use Cases

Convnets are used in computer vision applications. Image classification is one such problem: the input is an image, and the task is to assign that image to a category. Image-classification problems with small training datasets are identified as the most common use case when the organization is not a large technology company.

Framing a visual task

A computer receives an image and must assign it to a category using a relatively small training dataset. How should this task be described?

Identify the data: The input is visual data: an image.

Identify the task: The computer must assign the image to a category, so the task is image classification.

Identify the model family: A convolutional neural network, or convnet, is a deep-learning model used for this kind of computer vision work.

Identify the practical setting: The training dataset is small, which is an important use case described in the source.

This is a computer-vision image-classification problem using a convnet in a small-training-dataset setting.

Practice and Recall

EASY

A feature map has spatial dimensions of 26 × 26 and is processed with a 2 × 2 max-pooling window and stride 2. Explain the new spatial dimensions, describe what happens to the channels, and state what value each local window contributes to the output.

Hints
  • Apply the factor-of-two reduction separately to height and width.
  • Pooling changes spatial size but does not remove the channel dimension.
  • Each local window contributes its maximum value.
MEDIUM

A team wants to reduce spatial dimensions but is deciding between max pooling, average pooling, and a strided convolution. Explain the defining operation for each choice and why max pooling may preserve strong evidence of a local pattern.

Hints
  • Compare maximum, average, and convolutional processing.
  • Mention the role of stride for the convolutional option.
  • Relate the maximum to the strongest local presence of a pattern.

Key Takeaways

  • Max pooling replaces each local window with its maximum value, preserving the strongest local activation.
  • A 2 × 2 window with stride 2 reduces both spatial dimensions by a factor of 2, such as 26 × 26 becoming 13 × 13.
  • Pooling is performed independently for each channel, so spatial size changes while the channel dimension remains.
  • Downsampling reduces later processing cost and supports a hierarchy of features covering broader regions of the original input.
  • A convnet is a deep-learning model used in computer vision, including image classification with small training datasets.