Concepts / Choosing Network Architectures

Choosing Network Architectures

The hypothesis space is the set of possible transformations or solutions available to a learning algorithm.

  • Programming

The Architecture Question

When choosing a network architecture, it is tempting to ask only how many layers the network should have. A more useful question is: what kinds of input-output transformations can this architecture represent? That question leads to the idea of a hypothesis space. The hypothesis space is the set of possible transformations or solutions available to a learning algorithm. An architecture does not merely show which components are connected. Its topology determines the sequence of tensor operations that can be used to map inputs to outputs.

A network architecture sets the boundaries of what the network can represent. Learning then searches within those boundaries for particular weight values that produce a useful transformation.

What do you think happens?

Suppose a network receives an input tensor and has several layers, but every layer repeats the same restricted kind of linear operation. Does adding those layers automatically guarantee a larger hypothesis space?

  • Yes. Any additional layer always creates a larger hypothesis space.
  • No. Repeated linear layers can remain limited to linear operations.
  • The number of layers alone determines the answer.
Reveal answer

Answer: No. Repeated linear layers can remain limited to linear operations.

The kind of operation performed by the layers matters. More layers do not automatically add new kinds of transformations when the layers repeat a restricted operation.

Topology Sets the Operations

A topology specifies a series of tensor operations. Those operations define the family of transformations that the network can make available. For example, an architecture may arrange operations in a sequence, so the output of one operation becomes the input to the next. The topology therefore constrains how an input tensor can be changed as it moves through the network. The learned weight tensors do not change the basic operation sequence chosen by the topology. Instead, learning selects values for those weights and thereby selects one particular mapping from the available family.

flows toflows toproducesparameterizesparameterizesInput tensorxTensor operation Atopology-selectedTensor operation Btopology-selectedOutput tensoryWeight tensorslearned values
How does a chosen network architecture determine which tensor operations can connect the input to the output?

Read the diagram in two directions. From left to right, the topology describes the available route from an input tensor to an output tensor. From the weight-tensor branch, learning supplies values to the operations along that route. The route determines what kind of mapping is available; the learned values determine which particular mapping is selected.

A Hypothesis Space of Mappings

Hypothesis space: the set of possible transformations or solutions available to a learning algorithm.

For a network, it is useful to imagine the hypothesis space as a collection of possible input-output mappings. The architecture determines which kinds of mappings belong to that collection. The learning algorithm does not search over every imaginable transformation. It searches over the transformations permitted by the network's topology, using different learned weight values to select among them.

definescontainscontainscontainscan selectcan selectcan selectNetwork topologyoperation sequenceHypothesis spaceavailable mappingsMapping Ainput to outputMapping Binput to outputWeight valuesselected during learningMapping Cinput to output
What transformations are included in the hypothesis space, and how do they map inputs to possible outputs?

A solution can be absent because the architecture excludes it. If it is included, failure can still occur because parameter learning does not find it.

Tracing Two Architectures

Comparing a restricted sequence with a deeper restricted sequence

Consider two network designs. Design A applies one restricted linear operation. Design B applies several layers, but each layer repeats the same restricted linear kind of operation. What can you predict about their hypothesis spaces?

Identify the operation in Design A: Design A permits the transformation produced by its restricted linear operation, with the particular result selected by its weight values.

Trace the added layers in Design B: Design B adds depth, but its layers repeat the same restricted linear kind of operation. The added depth does not automatically introduce a new kind of operation.

Compare the available families: Because repeated linear layers remain limited to linear operations, depth alone does not guarantee a larger hypothesis space. The important question is whether the added layers change the kinds of operations available.

Separate architecture from learning: Even after deciding what transformations the architecture can represent, one more question remains: whether parameter learning will successfully find useful weight values within that available family.

The deeper design does not necessarily have a larger hypothesis space. Repeated restricted operations can leave the representational space restricted, while learning still has the separate task of finding a suitable included solution.

permitsrepeatsbelongs tocan remain withinDesign Aone restricted operationLinear mappingavailable transformationLineartransformationsrestricted familyDesign Bseveral repeated layersLinear mappingrepeated operation
Why can adding layers leave the set of representable input-output transformations unchanged or even restrict it?

The comparison is not claiming that every shallow and deep design has exactly the same hypothesis space. It shows why depth by itself is not enough to predict the result. The kind of operation performed by each layer matters. If added layers repeat a restricted operation, the resulting space can remain restricted rather than gaining the additional representational kinds that motivated the extra depth.

Representation Before Discovery

There are two separate reasons a network may fail to produce a desired solution. First, the solution may be excluded by the architecture: it does not belong to the network's hypothesis space. Second, the solution may belong to the hypothesis space, but parameter learning may fail to find it. These cases can look similar from the outside because both produce an unsuccessful result, but they require different diagnoses.

may bemay belearning findslearning missesDesired solutiontarget mappingExcludednot in hypothesis spaceFound solutionlearning succeedsIncludedin hypothesis spaceNot foundlearning does not select it
What is the difference between a solution being available inside the hypothesis space and the learning algorithm actually finding it?

Training as a Search

Once topology has defined the hypothesis space, training searches through that available space by learning values for the weight tensors. The search does not redesign the topology while it runs. It selects a particular transformation from the family that the topology made available. This gives a useful order for reasoning: identify the operations allowed by the architecture, determine the resulting hypothesis space, and then consider whether parameter learning can find a good member of that space.

constrainssets search regionchoosesNetwork topologyoperation sequenceHypothesis spacepossible transformationsLearn weight valuesparameter searchSelected mappingone available solution
How does training move from many possible transformations toward one selected solution?

When evaluating an architecture, describe both the operation sequence and the learning process. Saying only that the network has more layers leaves out the information needed to predict its hypothesis space.

Architecture Boundaries

Different network topologies create different boundaries around the transformations that an algorithm can represent. A topology with one sequence of tensor operations may make one family of mappings available, while another topology may make a different family available. The boundary is determined by the operations encoded in the architecture, not by layer count considered in isolation.

definesdefinesTopology Aoperation family AMapping family Aspace allowed by ATopology Boperation family BMapping family Bspace allowed by B
How do different network topologies create different boundaries around the transformations the algorithm can represent?

Diagnosing an unsuccessful result

A network does not produce a desired transformation. What two architecture-and-learning questions should be asked?

Check representation: Ask whether the desired transformation belongs to the hypothesis space defined by the network topology.

Check discovery: If the transformation is included, ask whether parameter learning successfully found the weight values that select it.

Interpret the result: If the transformation is excluded, changing the learning search alone cannot make that transformation available. If it is included but not found, the issue concerns the learning process rather than representational exclusion.

An unsuccessful result must be separated into a representational question and a parameter-learning question.

Common Reasoning Errors

  • Assuming that more layers always create a larger hypothesis space.

    Repeated linear layers remain limited to linear operations, so depth alone does not guarantee additional kinds of transformations.

    Fix: Inspect the kind of operation added by each layer, not just the number of layers.

  • Treating topology and learned weights as the same thing.

    Topology specifies the sequence of tensor operations, while learning selects values for the weight tensors involved in those operations.

    Fix: Separate the available family of mappings from the particular mapping selected by learned values.

  • Concluding that a solution is unrepresentable because training did not find it.

    The desired solution may be included in the hypothesis space even though parameter learning failed to find it.

    Fix: First test the representation question, then analyze the discovery question.

  • Assuming that a simple solution will automatically be learned if it exists.

    Whether a solution belongs to the hypothesis space and whether parameter learning finds it are different questions.

    Fix: State separately that the solution is available and that learning has or has not located it.

Architecture Check

MEDIUM

A proposed network has additional layers, but each added layer repeats the same restricted kind of operation. Explain why the phrase more layers is insufficient to establish that the network has a larger hypothesis space. Then describe the two questions you would ask if the network fails to produce a desired solution.

Hints
  • Start by identifying what the repeated operation allows.
  • Distinguish the family of transformations defined by topology from the weight values learned during training.
  • Use the words included and found to separate representation from discovery.
  1. A strong answer should say that repeated restricted operations can leave the hypothesis space restricted even when depth increases. It should then ask whether the desired solution belongs to the topology-defined hypothesis space and, if it does, whether parameter learning successfully finds it.

Key Takeaways

  • The hypothesis space is the set of possible transformations or solutions available to a learning algorithm.
  • Network topology determines the sequence of tensor operations and therefore constrains the transformations the network can represent.
  • Adding layers does not automatically enlarge the hypothesis space; repeated restricted operations can keep the space restricted.
  • A solution can be excluded from the hypothesis space, or it can be included but not found by parameter learning.
  • Architecture choice concerns representation, while training concerns selecting a particular mapping through learned weight values.