Concepts / Halfspaces

Halfspaces

The base class B provides the individual hypotheses used as building blocks.

  • Programming

From Weak Learners to a Larger Class

A weak learner begins with a restricted collection of possible hypotheses. This collection is called the base hypothesis class B. For example, B might be the class of decision stumps. An empirical risk minimization rule can choose one weak hypothesis from B, but boosting does not usually stop with that one hypothesis. It combines several weak hypotheses into a larger model.

The key separation is between components and the composed model. B supplies the individual hypotheses. L(B,T) describes the larger family obtained by selecting T hypotheses from B and combining their predictions with a halfspace controlled by a vector w in R^T. Thus, a member of L(B,T) is not simply one base hypothesis. It is a construction built from several members of B.

selectselectcombinecombinecontrolsBbase classh1member of BL(B,T)composed hypotheseshTmember of Bwvector in R^T
How are base hypotheses and their combination parameters related when forming a member of L(B,T)?

Following an Input Through the Construction

To understand the construction, follow one instance x. Each selected base hypothesis is applied to x. If the selected hypotheses are h1 through hT, their outputs are collected into the vector ψ(x) = (h1(x), ..., hT(x)) in R^T. The final hypothesis applies its halfspace to this prediction vector, not directly to the original instance x.

applyapplycollectcollectclassifyassignxinput instanceh1(x)predictionψ(x)(h1(x), ..., hT(x))halfspacecontrolled by wbinary labelpositive or negativehT(x)prediction
How does an input x become the vector of base-hypothesis predictions used by the final hypothesis?

Tracing One Instance

Suppose three selected base hypotheses are applied to an input x and produce the predictions h1(x), h2(x), and h3(x). What representation does the final halfspace receive?

Apply the base hypotheses: Each of the three selected hypotheses evaluates the same input x.

Collect the predictions: The three outputs are placed into one vector: ψ(x) = (h1(x), h2(x), h3(x)).

Apply the final rule: The halfspace controlled by w uses this vector as its input and produces the binary classification.

The final hypothesis operates on the vector of base-hypothesis predictions, not directly on the original input representation.

Why AdaBoost Fits L(B,T)

AdaBoost repeatedly obtains weak hypotheses from the base learning process. After T rounds, view the selected weak hypotheses as h1 through hT. AdaBoost then produces a composition of a halfspace over their predictions. This has exactly the structure used to define a member of L(B,T): T hypotheses come from B, and their prediction vector is combined using the vector w.

The membership argument can be checked as a sequence: AdaBoost selects weak hypotheses from B; the selected hypotheses are named h1 through hT; each input is transformed into ψ(x) = (h1(x), ..., hT(x)); a halfspace controlled by w is applied to ψ(x); the resulting composition is therefore a member of L(B,T).

When deciding whether a final model belongs to L(B,T), check its construction rather than asking whether it is itself a base hypothesis. The individual components belong to B, while the composed result belongs to L(B,T).

Halfspaces as Binary Decision Boundaries

A halfspace hypothesis assigns one of two labels to an input vector. Its decision is controlled by a weight vector and a bias parameter. Together, these parameters determine a separating hyperplane and the labels associated with the two sides of that boundary.

The weight vector controls the orientation of the separating boundary. The bias parameter also affects the position of the boundary. In two-dimensional feature space, the hyperplane is a line perpendicular to the weight vector. The line separates the plane into two regions: one side is associated with the positive label and the other with the negative label.

parameters defineorientspositionsseparatesseparatesfeature planeno separating line shownhyperplaneline in two dimensionspositive regionone sidenegative regionother sideweight vectorperpendicular directionbiasboundary position
How do the weight vector and bias define a boundary whose two sides correspond to positive and negative labels?

Imagine a two-dimensional feature space in which every instance must receive one of two labels. A halfspace places a line through that space. Instances on the side described as above the line receive the positive label, while instances on the side described as below the line receive the negative label. These descriptions are geometric: above and below depend on the line and the direction of the weight vector, not necessarily on physical up and down.

Common Mistakes

  • Treating B as the final boosted hypothesis class

    B supplies individual building blocks. The composed result is represented in the larger class L(B,T).

    Fix: Distinguish the source of the components, B, from the class of their halfspace combinations, L(B,T).

  • Applying the final halfspace directly to x

    The construction first applies each selected base hypothesis and collects the outputs into ψ(x) = (h1(x), ..., hT(x)).

    Fix: Trace the two stages: x becomes ψ(x), then the halfspace acts on ψ(x).

  • Confusing the boundary with a classified region

    The hyperplane is the separating boundary. The regions on its two sides are associated with the two labels.

    Fix: Name the line as the boundary and the two sides as the positive and negative regions.

  • Interpreting above and below as universal physical directions

    The meaning of above and below depends on the geometric relationship to the hyperplane and the direction of the weight vector.

    Fix: Describe the labels relative to the separating boundary and its orientation.

Check Your Understanding

MEDIUM

An AdaBoost construction has selected T weak hypotheses from B. Explain, in order, what happens to an input x before the final binary label is produced. Then explain what the weight vector and bias contribute to the halfspace stage.

Hints
  • Start by naming the outputs h1(x) through hT(x).
  • Use the notation ψ(x) for the collected prediction vector.
  • For the geometric part, distinguish the boundary from the two regions on its sides.

Answer Structure

What argument shows that the AdaBoost output belongs to L(B,T)?

Identify the components: The weak hypotheses produced across the rounds are viewed as h1 through hT, and each comes from B.

Identify the representation: For an input x, their outputs form ψ(x) = (h1(x), ..., hT(x)).

Identify the final operation: A halfspace controlled by w is applied to ψ(x).

Match the definition: This is precisely the structure of a hypothesis in L(B,T): T base hypotheses combined through a halfspace.

The AdaBoost output belongs to L(B,T) because its construction matches the definition of that composed hypothesis class.

Key Takeaways

  1. B is the base hypothesis class that supplies individual weak hypotheses.
  2. L(B,T) is formed by selecting T hypotheses from B and combining their predictions with a halfspace controlled by w in R^T.
  3. An input x is transformed into ψ(x) = (h1(x), ..., hT(x)) before the final halfspace makes its decision.
  4. AdaBoost belongs to L(B,T) because its output has this same composition structure.
  5. For binary classification, a halfspace uses a weight vector and bias to define a separating hyperplane; its two sides correspond to positive and negative labels.

Key Takeaways

  • The base class B provides the individual hypotheses used as building blocks.
  • The class L(B,T) combines T such hypotheses through a halfspace controlled by a vector w.
  • The final classifier receives the vector of base-hypothesis predictions, ψ(x), rather than directly applying the halfspace to x.
  • AdaBoost's output belongs to L(B,T) because it has exactly this composition structure.
  • A halfspace separates two labels with a hyperplane whose orientation and position are determined by its parameters.