Subsampling layers
Deep convolutional networks organize processing through convolutional, subsampling, and fully connected layers.
From Image to Recognition
A deep convolutional network does not treat handwritten-character recognition as one indivisible operation. It organizes processing as a sequence. Convolutional layers and subsampling layers alternate, and the resulting representations are then passed to fully connected final layers. This sequence gives the network a route from local information in an input image to recognition of the handwritten character.
A First-Layer Feature Map
A feature map is a collection of responses produced by units at different locations. Those units apply shared weights at shifted locations, so the map records where the same learned operation responds across the incoming array. The operation is not separately learned for every position. Instead, units in the same feature map use shared weights.
Reading the first convolutional layer
Interpret the source network's first-layer feature maps.
Identify the maps: The first layer has 6 feature maps.
Identify the units: Each map has 28 × 28 units, so each map is an organized collection of responses across locations.
Identify each local view: Each unit uses a 5 × 5 receptive field from the incoming array.
Identify the learned parameters: Each map has 25 adjustable weights associated with its shared operation.
The first layer produces six spatially organized activity patterns. Within any one map, different units apply the same learned weights at shifted locations.
Receptive Fields and Overlap
A receptive field is the local input region seen by one unit. It specifies which part of the incoming array can contribute to that unit's response.
In the source network, first-layer units have 5 × 5 receptive fields. Neighboring units can view regions that overlap. Specifically, the first-layer fields overlap by four columns and five rows. This overlap describes how neighboring units draw their local views from the incoming array. It does not mean that each unit has its own separate set of weights: units in the same feature map still use shared weights.
A feature-map value has a unit-level explanation and a map-level explanation. At the unit level, its receptive field identifies the local input region. At the map level, its operation is related to the responses at other locations because the weights are shared.
What Subsampling Changes
Subsampling layers are the stages that follow convolutional stages in the alternating architecture. Their role in the representation is to change the spatial resolution of feature maps while preserving useful detected features. The result is not a single number. It is a collection of activity patterns that continues through later convolutional and subsampling stages.
Building Character Evidence
The architecture can be understood as a gradual organization of information. Convolutional layers create local activity patterns by applying shared operations at different locations. Subsampling layers transform those patterns into later representations. Repeated convolutional and subsampling stages prepare the signal for fully connected final layers, which receive the resulting representations. In this way, the network provides a route from local input information to the recognition of a handwritten character.
Tracing one recognition pathway
Trace how information can move through the architecture when the input is a handwritten character.
Start with local information: Units in an early convolutional layer view local input regions through their receptive fields.
Form feature-map activity: Shared weights are applied at shifted locations, creating activity patterns across feature maps.
Pass through subsampling: The activity patterns move through subsampling stages as the network organizes the representation.
Repeat the preparation: Convolutional and subsampling layers alternate, preparing the signal for later processing.
Reach the final layers: Fully connected final layers receive the representations produced by the earlier stages.
The network moves from local views and shared feature responses toward representations that support handwritten-character recognition.
Mistakes to Avoid
Treating a feature map as if every location learned a different operation.
Units in the same feature map apply shared weights at shifted locations.
Fix:
Describe the map as responses from one shared learned operation applied across locations.Confusing a receptive field with the whole feature map.
A receptive field specifies the local input region seen by one unit.
Fix:
Use the unit-level view: one first-layer unit has a 5 × 5 receptive field in the source network.Assuming overlapping receptive fields require separate weights.
Overlap concerns the local input regions, while weight sharing concerns the learned operation.
Fix:
Explain both facts together: neighboring fields may overlap, and units in the same map still share weights.Leaving subsampling out of the network sequence.
The described network alternates convolutional and subsampling layers before the fully connected final layers.
Fix:
State the sequence at the network level before explaining individual units.Explaining recognition as a single feature-map response.
The network produces collections of activity patterns and passes their resulting representations to fully connected final layers.
Fix:
Trace the representation across stages from local information to the final layers.
Check Your Understanding
Explain the first layer of the source network in three levels of description: the network-level sequence, the feature-map-level operation, and the unit-level receptive field.
Hints
- At the network level, name the layer types that alternate before the fully connected final layers.
- At the feature-map level, explain what is shared across locations.
- At the unit level, include the first-layer receptive-field size and explain how neighboring fields can overlap.
A learner says, “The six first-layer feature maps each contain unrelated responses because every location has separate weights.” Correct the statement using the source network's 28 × 28 maps, 5 × 5 receptive fields, and shared weights.
Hints
- Separate the number of feature maps from the number of locations within one map.
- Explain what the 25 adjustable weights per map represent.
- Mention that receptive-field overlap does not remove weight sharing.
Key Takeaways
- A deep convolutional network organizes processing through convolutional, subsampling, and fully connected layers.
- A feature map collects responses from units that apply shared weights at shifted locations.
- A receptive field is the local input region seen by one unit.
- Neighboring receptive fields can overlap; in the source network's first layer, the 5 × 5 fields overlap by four columns and five rows.
- Alternating convolutional and subsampling stages organize local information into representations received by fully connected final layers for handwritten-character recognition.
Key Takeaways
- Convolutional layers create feature-map activity by applying shared weights at different locations.
- Subsampling layers change the spatial representation and pass organized activity patterns to later stages.
- A receptive field describes the local input region seen by one unit, and neighboring fields may overlap.
- The architecture alternates convolutional and subsampling layers before fully connected final layers use the resulting representations for handwritten-character recognition.