Concepts / Introduction to Convolutional Neural Networks

Introduction to Convolutional Neural Networks

Convolution focuses on local patches rather than treating the entire image as one undivided pattern.

  • Programming
Interactive lab

Try it: Convolution

How a convolution layer slides a small kernel over an image, multiplying and summing at each position to build a feature map.

How it works

  1. Place the 3×3 kernel over the top-left 3×3 window of the input.
  2. Multiply each input value by the kernel weight on top of it and add the nine products — that sum is one output cell.
  3. Slide the window by the stride and repeat until every position is covered.
  4. Zero padding adds a border so edge pixels are covered too; output size = ⌊(n + 2p − k) / s⌋ + 1.

Default run (11 steps): 5×5 input, 3×3 kernel, stride 1. Output is ⌊(5 + 2×0 − 3) / 1⌋ + 1 = 3 per side. … Done: the kernel visited all 9 positions and produced a 3×3 feature map.

Simplified: One 5×5 single-channel input and one 3×3 kernel. As in deep-learning libraries, the kernel is not flipped (strictly, cross-correlation).

Educational simulation

Loading the simulation…

From Whole Images to Local Patches

A convolutional neural network does not learn from an image only as one enormous, undivided object. Instead, a convolution layer examines small local regions, called patches, of the input. For an image, each patch is a small two-dimensional window. This local view allows the layer to learn patterns found in particular regions rather than requiring every learned pattern to involve every pixel at once.

examinedWhole imageLocal featureLocal patch
What portion of the input does a convolution inspect at one time, and how does this differ from treating the entire image as one undivided pattern?

Tracing One Transformation

A convolution operation applies a learned transformation to one local patch and then reuses that same transformation at other locations in the input. The transformation is not redesigned separately for every patch. As the operation considers different local regions, it produces values that record how strongly the learned transformation responds at those locations.

extractapplyreusereuseproduceInput feature mapPatch 1FilterResponsesPatch 2Patch 3
How does the same filter move across an input feature map and apply the same calculation to each local patch?

Following a learned pattern

Imagine that a filter encodes one particular aspect of the input. How does the convolution layer use it across an image?

Select a patch: The layer examines one small local region of the input feature map.

Apply the filter: The layer applies its learned transformation to that patch.

Move to another location: The layer examines another local patch while reusing the same transformation.

Record the response: The output records how strongly the learned transformation responds at each considered location.

The layer searches for the same learned pattern across multiple local regions instead of learning a separate transformation for every location.

Feature Map Dimensions

A feature map is a three-dimensional tensor with height, width, and depth. Height and width describe spatial position. Depth, also called the channels axis, represents channels.

The height and width axes tell us where a feature appears spatially. The depth axis holds channels associated with that spatial structure. In an input image, channels can represent information such as RGB color channels. In an output feature map, channels represent filters and the responses produced by those filters; they no longer represent specific RGB colors.

axisaxisaxisFeature mapthree-dimensional tensorHeightspatial positionWidthspatial positionDepthchannels
Where are the height, width, and depth axes in a feature map, and what does each axis describe?
AxisMeaning
HeightSpatial position along one image dimension
WidthSpatial position along the other image dimension
DepthChannels of the feature map

The roles of the three feature-map axes

Filters and Response Maps

A filter encodes a specific aspect of the input data. When a convolution layer uses several filters, each filter produces one channel in the output feature map. That channel is a response map: a two-dimensional grid showing the response of that filter at different locations in the input.

inspectinspectproduceproducechannelchannelInput feature mapFilter AResponse map AOutput feature mapFilter BResponse map B
How do the values produced by a filter at different patch positions become a new response map in the output feature map?

Reading an output channel

A convolution layer uses several filters. What does one output channel represent?

Choose one filter: The filter encodes one specific aspect of the input data.

Consider its locations: The same filter responds at different local patch positions in the input.

Arrange the responses: The responses form a two-dimensional grid whose positions correspond to locations in the input.

Interpret the channel: That grid is a response map and becomes one channel of the output feature map.

One filter produces one response-map channel, while several filters produce several output channels.

Recognizing Patterns Across Locations

Convolutional layers learn translation-invariant patterns. If a network learns a pattern in one location, such as the lower-right corner of a picture, it can recognize that pattern elsewhere, such as the upper-left corner. The reason is that the same local transformation is applied at different locations. The pattern does not have to be learned as an entirely new pattern for every position.

located atrecognized bylocated atrecognized byLearned patternLower-rightPattern responseLearned patternUpper-leftPattern response
What changes in the response map when a learned pattern shifts to a different location in the input, and what remains recognizable?

Building Spatial Hierarchies

Convolutional networks can learn spatial hierarchies of patterns. Early convolution layers can learn small local patterns such as edges. Later layers can learn larger patterns made from features learned by earlier layers. Further layers can continue combining those patterns into increasingly complex and abstract visual concepts.

combinecombineEdgessmall local patternsLarger patternscombined featuresComplex conceptsincreasingly abstract
How do patterns detected in earlier feature maps combine into larger and more complex patterns in deeper layers?
StagePatterns emphasizedRole
Earlier layersSmall local patterns such as edgesProvide features that later layers can use as building blocks
Later layersLarger patterns made from earlier featuresCombine earlier features into more complex and abstract visual concepts

When tracing a convolutional network, ask two questions at each stage: what local or combined pattern is being represented, and how can the next stage use it as a building block? This keeps the progression from small local features to larger abstract concepts clear.

Common Mistakes

  • Treating convolution as if it learned one pattern from the entire image.

    A convolution layer examines small local regions or patches of the input.

    Fix: Describe the operation as applying a learned transformation to local patches across the input.

  • Assuming that every patch receives a newly designed transformation.

    The same transformation is reused across patches.

    Fix: Explain that one learned transformation is applied at different locations to search for the same pattern.

  • Confusing output depth with RGB color channels.

    In an output feature map, channels represent filters and the responses produced by those filters.

    Fix: Interpret each output channel as a response map associated with a filter.

  • Thinking that translation invariance means the response map loses all location information.

    A response map shows the response of a filter at different input locations.

    Fix: Explain that the same pattern can be recognized at different locations while the response map records those locations.

  • Expecting the earliest layer to represent the most abstract concept.

    The progression begins with small local patterns, which later layers combine into larger and more abstract concepts.

    Fix: Trace the hierarchy from local features such as edges to increasingly complex visual concepts.

Practice: Trace the Data

MEDIUM

A convolution layer examines local patches using one learned filter and then uses several filters in the layer. Describe what one filter contributes to the output feature map. In your answer, identify the input region it examines, explain why the same transformation can be used at multiple locations, and name the output structure created by that filter.

Hints
  • Start with the definition of a local patch.
  • Remember that the transformation is reused rather than redesigned for each patch.
  • A single filter produces one two-dimensional response map, which becomes one output channel.

Practice answer

What does one filter contribute to an output feature map?

Input region: The filter is applied to local patches of the input feature map.

Shared transformation: The same learned transformation is reused at different patch locations.

Response grid: The filter produces responses at those locations, forming a two-dimensional response map.

Output channel: That response map becomes one channel in the three-dimensional output feature map.

One filter searches for one learned aspect of the input across multiple locations and contributes one response-map channel to the output.

Key Takeaways

  1. Convolution examines small local patches instead of treating the entire image as one undivided pattern.
  2. The same learned transformation is reused across different patches and locations.
  3. A feature map has height and width spatial axes plus a depth axis for channels.
  4. Each filter produces a two-dimensional response map that becomes one output channel.
  5. Convolutional networks recognize patterns across locations and build spatial hierarchies from small local features to larger, more abstract concepts.

Key Takeaways

  • Convolution focuses on local patches rather than the whole image as one undivided pattern.
  • A shared transformation searches for the same learned pattern at multiple locations.
  • Feature maps use height and width for spatial position and depth for channels.
  • Filters create response maps, which become channels in the output feature map.
  • Deeper layers combine earlier local features into increasingly complex and abstract visual concepts.