Concepts / Representation Learning with Multiple Layers

Representation Learning with Multiple Layers

The hypothesis space is the set of possible transformations or solutions available to a learning algorithm.

  • Programming

The Search Space Before Learning

When a learning system fails to produce a desired solution, there are two different possibilities. The solution may not be available within the network's permitted transformations, or the solution may be available but the learning process may fail to find it. Representation learning with multiple layers becomes clearer when these possibilities are kept separate.

A hypothesis space is the set of possible transformations or solutions available to a learning algorithm.

containsselectsHypothesis spaceavailable transformationsTransformation familyallowed by topologySelected mappingchosen by learned weights
Which transformations can the learning algorithm search through?

Tracing a Topology

A network topology is more than a drawing of connected layers. It specifies a series of tensor operations. An input tensor is passed through the operations made available by the arrangement of layers and connections, producing an output tensor. The topology therefore constrains the kinds of input-output mappings the network can express.

tensor entersresult passestensor exitsInput tensorOperation 1allowed by topologyOperation 2allowed by topologyOutput tensor
How does the arrangement of layers and connections determine the sequence of tensor operations?

Following One Possible Mapping

Consider a network topology that provides two tensor operations in sequence. What determines the family of input-output mappings that this network can provide?

Identify the topology: The arrangement of layers and connections specifies that the input passes through operation 1 and then operation 2.

Identify the available family: The sequence of operations, rather than the diagram alone, determines the kinds of transformations available to the network.

Identify the learned choice: Learning searches for values of the weight tensors involved in those operations. Different learned values can select different mappings from the family allowed by the topology.

Topology determines what kinds of mappings are available; learned weights select a particular mapping from that available family.

Depth and Operation Type

Adding a layer does not automatically enlarge the hypothesis space. The important question is what operation the added layer performs. If the added layers repeat the same restricted kind of operation, the resulting family of functions can remain restricted. In particular, repeated linear layers remain limited to linear operations. Therefore, depth by itself does not determine the full range of functions a network can represent.

allowsallowsOne linear layerlinear operationsLinear mappingsavailable familyRepeated linearlayerslinear operationsLinear mappingsstill restricted
When layers are added, do they create new possible input-output functions, or can they re-express transformations already available?

What do you think happens?

A network receives an additional layer, but the added layer repeats the same restricted linear operation. Should you conclude automatically that the hypothesis space has become larger?

  • Yes, because every additional layer always adds new functions
  • No, because repeated linear layers can remain limited to linear operations
  • The topology has no effect on the hypothesis space
Reveal answer

Answer: No, because repeated linear layers can remain limited to linear operations.

The type of operation matters more than layer count alone. Extra depth that repeats a restricted operation may leave the available family restricted.

Representable versus Found

A solution being included in the hypothesis space means that the topology and its permitted operations can represent that solution. It does not mean that training will discover it. Once the topology specifies an available family, parameter learning searches for a good set of values for the weight tensors. That search can fail even when the desired solution belongs to the family.

is tested againstcan includecan excludeis searched bymay producemay fail to findDesired solutionTopologydefines available mappingsIncluded solutionrepresentableParameter learningsearches weight valuesFound solutionlearning succeedsExcluded solutionnot representableMissed solutionlearning fails
What is the difference between a solution being representable by the network and the learning algorithm actually finding it?

Separating Representation from Training

Suppose a desired transformation is compatible with the operations specified by a network topology, but training does not produce it. Which question has a negative answer: whether the solution is available, or whether learning found it?

Check availability: Because the desired transformation is compatible with the topology, it belongs to the network's hypothesis space.

Check the learning result: Training did not produce the desired transformation, so parameter learning failed to find an included solution.

Separate the diagnosis: This is different from representational exclusion. In representational exclusion, the topology would not permit the desired transformation in the first place.

A solution can be representable but still not be found by the learning process.

QuestionWhat it diagnosesMeaning of failure
Can the topology express the desired transformation?Representational capacityThe solution is excluded from the hypothesis space
Can parameter learning select the desired transformation?Learning successThe solution may be included but was not found

Representation and successful learning are separate questions.

Common Reasoning Errors

  • Assuming that every added layer enlarges the hypothesis space.

    If the added layers repeat the same restricted kind of operation, such as a linear operation, the resulting hypothesis space can remain restricted.

    Fix: Inspect the operations introduced by the topology, not only the number of layers.

  • Treating a network diagram as separate from the transformations it permits.

    Topology determines the sequence of tensor operations available for mapping inputs to outputs.

    Fix: Trace the input through the operations specified by the arrangement of layers and connections.

  • Concluding that a failed training result proves the solution is impossible for the network.

    A desired solution can belong to the hypothesis space while parameter learning fails to find it.

    Fix: Ask separately whether the solution is representable and whether the learning process found it.

  • Assuming that a simple solution will automatically be learned.

    The existence of a simple solution does not answer the separate question of whether the training setup will find it.

    Fix: Keep hypothesis-space inclusion and successful parameter learning as distinct claims.

Topology Audit Practice

MEDIUM

A proposed network has additional layers, but every added layer repeats the same restricted operation already used earlier. Explain whether the layer count alone proves that the hypothesis space is larger. Then describe the two separate questions you would ask if training fails to produce a desired transformation.

Hints
  • Start by identifying the operation performed by the added layers.
  • Use the distinction between topology-defined availability and parameter-learning success.
  • Remember that repeated linear layers remain limited to linear operations.

When analyzing a multilayer network, use this order: identify the topology, trace the sequence of tensor operations, describe the resulting hypothesis space, and only then ask whether parameter learning can find the desired mapping.

Key Takeaways

  1. The hypothesis space is the set of transformations or solutions available to a learning algorithm.
  2. Network topology determines the sequence of tensor operations and therefore constrains the available mappings.
  3. Adding layers does not necessarily enlarge the hypothesis space; repeated restricted operations can leave it restricted.
  4. A desired solution can be excluded by the topology, or it can be included but missed by parameter learning.
  5. Layer count is less informative than the kind of operation performed by the layers.

Key Takeaways

  • A hypothesis space describes the possible transformations a learning algorithm can search through.
  • Topology defines the available sequence of tensor operations, while learned weights select a particular mapping from that family.
  • More layers do not automatically mean a larger hypothesis space, especially when they repeat restricted operations such as linear operations.
  • Failure can mean either that the topology excludes the desired solution or that learning failed to find a solution that was available.