Choosing Network Architectures
The hypothesis space is the set of possible transformations or solutions available to a learning algorithm.
The Architecture Question
When choosing a network architecture, it is tempting to ask only how many layers the network should have. A more useful question is: what kinds of input-output transformations can this architecture represent? That question leads to the idea of a hypothesis space. The hypothesis space is the set of possible transformations or solutions available to a learning algorithm. An architecture does not merely show which components are connected. Its topology determines the sequence of tensor operations that can be used to map inputs to outputs.
A network architecture sets the boundaries of what the network can represent. Learning then searches within those boundaries for particular weight values that produce a useful transformation.
What do you think happens?
Suppose a network receives an input tensor and has several layers, but every layer repeats the same restricted kind of linear operation. Does adding those layers automatically guarantee a larger hypothesis space?
Reveal answer
Answer: No. Repeated linear layers can remain limited to linear operations.
The kind of operation performed by the layers matters. More layers do not automatically add new kinds of transformations when the layers repeat a restricted operation.
Topology Sets the Operations
A topology specifies a series of tensor operations. Those operations define the family of transformations that the network can make available. For example, an architecture may arrange operations in a sequence, so the output of one operation becomes the input to the next. The topology therefore constrains how an input tensor can be changed as it moves through the network. The learned weight tensors do not change the basic operation sequence chosen by the topology. Instead, learning selects values for those weights and thereby selects one particular mapping from the available family.
Read the diagram in two directions. From left to right, the topology describes the available route from an input tensor to an output tensor. From the weight-tensor branch, learning supplies values to the operations along that route. The route determines what kind of mapping is available; the learned values determine which particular mapping is selected.
A Hypothesis Space of Mappings
Hypothesis space: the set of possible transformations or solutions available to a learning algorithm.
For a network, it is useful to imagine the hypothesis space as a collection of possible input-output mappings. The architecture determines which kinds of mappings belong to that collection. The learning algorithm does not search over every imaginable transformation. It searches over the transformations permitted by the network's topology, using different learned weight values to select among them.
A solution can be absent because the architecture excludes it. If it is included, failure can still occur because parameter learning does not find it.
Tracing Two Architectures
Comparing a restricted sequence with a deeper restricted sequence
Consider two network designs. Design A applies one restricted linear operation. Design B applies several layers, but each layer repeats the same restricted linear kind of operation. What can you predict about their hypothesis spaces?
Identify the operation in Design A: Design A permits the transformation produced by its restricted linear operation, with the particular result selected by its weight values.
Trace the added layers in Design B: Design B adds depth, but its layers repeat the same restricted linear kind of operation. The added depth does not automatically introduce a new kind of operation.
Compare the available families: Because repeated linear layers remain limited to linear operations, depth alone does not guarantee a larger hypothesis space. The important question is whether the added layers change the kinds of operations available.
Separate architecture from learning: Even after deciding what transformations the architecture can represent, one more question remains: whether parameter learning will successfully find useful weight values within that available family.
The deeper design does not necessarily have a larger hypothesis space. Repeated restricted operations can leave the representational space restricted, while learning still has the separate task of finding a suitable included solution.
The comparison is not claiming that every shallow and deep design has exactly the same hypothesis space. It shows why depth by itself is not enough to predict the result. The kind of operation performed by each layer matters. If added layers repeat a restricted operation, the resulting space can remain restricted rather than gaining the additional representational kinds that motivated the extra depth.
Representation Before Discovery
There are two separate reasons a network may fail to produce a desired solution. First, the solution may be excluded by the architecture: it does not belong to the network's hypothesis space. Second, the solution may belong to the hypothesis space, but parameter learning may fail to find it. These cases can look similar from the outside because both produce an unsuccessful result, but they require different diagnoses.
Training as a Search
Once topology has defined the hypothesis space, training searches through that available space by learning values for the weight tensors. The search does not redesign the topology while it runs. It selects a particular transformation from the family that the topology made available. This gives a useful order for reasoning: identify the operations allowed by the architecture, determine the resulting hypothesis space, and then consider whether parameter learning can find a good member of that space.
When evaluating an architecture, describe both the operation sequence and the learning process. Saying only that the network has more layers leaves out the information needed to predict its hypothesis space.
Architecture Boundaries
Different network topologies create different boundaries around the transformations that an algorithm can represent. A topology with one sequence of tensor operations may make one family of mappings available, while another topology may make a different family available. The boundary is determined by the operations encoded in the architecture, not by layer count considered in isolation.
Diagnosing an unsuccessful result
A network does not produce a desired transformation. What two architecture-and-learning questions should be asked?
Check representation: Ask whether the desired transformation belongs to the hypothesis space defined by the network topology.
Check discovery: If the transformation is included, ask whether parameter learning successfully found the weight values that select it.
Interpret the result: If the transformation is excluded, changing the learning search alone cannot make that transformation available. If it is included but not found, the issue concerns the learning process rather than representational exclusion.
An unsuccessful result must be separated into a representational question and a parameter-learning question.
Common Reasoning Errors
Assuming that more layers always create a larger hypothesis space.
Repeated linear layers remain limited to linear operations, so depth alone does not guarantee additional kinds of transformations.
Fix:
Inspect the kind of operation added by each layer, not just the number of layers.Treating topology and learned weights as the same thing.
Topology specifies the sequence of tensor operations, while learning selects values for the weight tensors involved in those operations.
Fix:
Separate the available family of mappings from the particular mapping selected by learned values.Concluding that a solution is unrepresentable because training did not find it.
The desired solution may be included in the hypothesis space even though parameter learning failed to find it.
Fix:
First test the representation question, then analyze the discovery question.Assuming that a simple solution will automatically be learned if it exists.
Whether a solution belongs to the hypothesis space and whether parameter learning finds it are different questions.
Fix:
State separately that the solution is available and that learning has or has not located it.
Architecture Check
A proposed network has additional layers, but each added layer repeats the same restricted kind of operation. Explain why the phrase more layers is insufficient to establish that the network has a larger hypothesis space. Then describe the two questions you would ask if the network fails to produce a desired solution.
Hints
- Start by identifying what the repeated operation allows.
- Distinguish the family of transformations defined by topology from the weight values learned during training.
- Use the words included and found to separate representation from discovery.
- A strong answer should say that repeated restricted operations can leave the hypothesis space restricted even when depth increases. It should then ask whether the desired solution belongs to the topology-defined hypothesis space and, if it does, whether parameter learning successfully finds it.
Key Takeaways
- The hypothesis space is the set of possible transformations or solutions available to a learning algorithm.
- Network topology determines the sequence of tensor operations and therefore constrains the transformations the network can represent.
- Adding layers does not automatically enlarge the hypothesis space; repeated restricted operations can keep the space restricted.
- A solution can be excluded from the hypothesis space, or it can be included but not found by parameter learning.
- Architecture choice concerns representation, while training concerns selecting a particular mapping through learned weight values.