Representation Learning with Multiple Layers
The hypothesis space is the set of possible transformations or solutions available to a learning algorithm.
The Search Space Before Learning
When a learning system fails to produce a desired solution, there are two different possibilities. The solution may not be available within the network's permitted transformations, or the solution may be available but the learning process may fail to find it. Representation learning with multiple layers becomes clearer when these possibilities are kept separate.
A hypothesis space is the set of possible transformations or solutions available to a learning algorithm.
Tracing a Topology
A network topology is more than a drawing of connected layers. It specifies a series of tensor operations. An input tensor is passed through the operations made available by the arrangement of layers and connections, producing an output tensor. The topology therefore constrains the kinds of input-output mappings the network can express.
Following One Possible Mapping
Consider a network topology that provides two tensor operations in sequence. What determines the family of input-output mappings that this network can provide?
Identify the topology: The arrangement of layers and connections specifies that the input passes through operation 1 and then operation 2.
Identify the available family: The sequence of operations, rather than the diagram alone, determines the kinds of transformations available to the network.
Identify the learned choice: Learning searches for values of the weight tensors involved in those operations. Different learned values can select different mappings from the family allowed by the topology.
Topology determines what kinds of mappings are available; learned weights select a particular mapping from that available family.
Depth and Operation Type
Adding a layer does not automatically enlarge the hypothesis space. The important question is what operation the added layer performs. If the added layers repeat the same restricted kind of operation, the resulting family of functions can remain restricted. In particular, repeated linear layers remain limited to linear operations. Therefore, depth by itself does not determine the full range of functions a network can represent.
What do you think happens?
A network receives an additional layer, but the added layer repeats the same restricted linear operation. Should you conclude automatically that the hypothesis space has become larger?
Reveal answer
Answer: No, because repeated linear layers can remain limited to linear operations.
The type of operation matters more than layer count alone. Extra depth that repeats a restricted operation may leave the available family restricted.
Representable versus Found
A solution being included in the hypothesis space means that the topology and its permitted operations can represent that solution. It does not mean that training will discover it. Once the topology specifies an available family, parameter learning searches for a good set of values for the weight tensors. That search can fail even when the desired solution belongs to the family.
Separating Representation from Training
Suppose a desired transformation is compatible with the operations specified by a network topology, but training does not produce it. Which question has a negative answer: whether the solution is available, or whether learning found it?
Check availability: Because the desired transformation is compatible with the topology, it belongs to the network's hypothesis space.
Check the learning result: Training did not produce the desired transformation, so parameter learning failed to find an included solution.
Separate the diagnosis: This is different from representational exclusion. In representational exclusion, the topology would not permit the desired transformation in the first place.
A solution can be representable but still not be found by the learning process.
| Question | What it diagnoses | Meaning of failure |
|---|---|---|
| Can the topology express the desired transformation? | Representational capacity | The solution is excluded from the hypothesis space |
| Can parameter learning select the desired transformation? | Learning success | The solution may be included but was not found |
Representation and successful learning are separate questions.
Common Reasoning Errors
Assuming that every added layer enlarges the hypothesis space.
If the added layers repeat the same restricted kind of operation, such as a linear operation, the resulting hypothesis space can remain restricted.
Fix:
Inspect the operations introduced by the topology, not only the number of layers.Treating a network diagram as separate from the transformations it permits.
Topology determines the sequence of tensor operations available for mapping inputs to outputs.
Fix:
Trace the input through the operations specified by the arrangement of layers and connections.Concluding that a failed training result proves the solution is impossible for the network.
A desired solution can belong to the hypothesis space while parameter learning fails to find it.
Fix:
Ask separately whether the solution is representable and whether the learning process found it.Assuming that a simple solution will automatically be learned.
The existence of a simple solution does not answer the separate question of whether the training setup will find it.
Fix:
Keep hypothesis-space inclusion and successful parameter learning as distinct claims.
Topology Audit Practice
A proposed network has additional layers, but every added layer repeats the same restricted operation already used earlier. Explain whether the layer count alone proves that the hypothesis space is larger. Then describe the two separate questions you would ask if training fails to produce a desired transformation.
Hints
- Start by identifying the operation performed by the added layers.
- Use the distinction between topology-defined availability and parameter-learning success.
- Remember that repeated linear layers remain limited to linear operations.
When analyzing a multilayer network, use this order: identify the topology, trace the sequence of tensor operations, describe the resulting hypothesis space, and only then ask whether parameter learning can find the desired mapping.
Key Takeaways
- The hypothesis space is the set of transformations or solutions available to a learning algorithm.
- Network topology determines the sequence of tensor operations and therefore constrains the available mappings.
- Adding layers does not necessarily enlarge the hypothesis space; repeated restricted operations can leave it restricted.
- A desired solution can be excluded by the topology, or it can be included but missed by parameter learning.
- Layer count is less informative than the kind of operation performed by the layers.
Key Takeaways
- A hypothesis space describes the possible transformations a learning algorithm can search through.
- Topology defines the available sequence of tensor operations, while learned weights select a particular mapping from that family.
- More layers do not automatically mean a larger hypothesis space, especially when they repeat restricted operations such as linear operations.
- Failure can mean either that the topology excludes the desired solution or that learning failed to find a solution that was available.