Concepts / Base Hypothesis Class

Base Hypothesis Class

The base class B provides the individual hypotheses used as building blocks.

  • Programming

From Weak Learners to a Larger Class

A weak learner does not begin with every possible hypothesis. It begins with a restricted collection of candidates. That collection is the base hypothesis class, written B. A weak learning rule can select one hypothesis from B, but boosting uses several selected hypotheses together rather than treating only one weak hypothesis as the final model.

containscontainscontainsbuilding blockbuilding blockbuilding blockBbase hypothesis classh1selected hypothesish2selected hypothesisL(B,T)larger hypothesis classhTselected hypothesis
What does the base hypothesis class contain, and how do its members become building blocks for a larger class?

Selecting T Members and w

The class L(B,T) separates the construction into two choices. First, select T hypotheses from B and name them h1 through hT. Second, choose a vector w in R^T. These selected hypotheses provide the individual predictions, while w controls the homogeneous halfspace applied to those predictions. Thus, one member of L(B,T) is determined by the selected base hypotheses together with the vector w.

selectselectselectprediction componentprediction componentprediction componentcontrols halfspaceBsource classh1member of Bh2member of BhTmember of Bwvector in R^TL(B,T)one composed hypothesis
How do the selected hypotheses and the vector w jointly define one member of L(B,T)?

Constructing One Member of L(B,T)

Suppose B supplies three selected hypotheses, h1, h2, and h3. Describe the information needed to define the corresponding composed hypothesis when T is 3.

Select the base hypotheses: Choose h1, h2, and h3 from B. These are the three individual hypotheses used as building blocks.

Choose the parameter vector: Choose a vector w in R^3. Its three coordinates correspond to the three selected hypotheses.

Form the composed hypothesis: For an input x, collect the three base-hypothesis outputs and apply the homogeneous halfspace defined by w to that prediction vector.

The selected hypotheses h1, h2, and h3 together with w define one member of L(B,3).

Following an Input Through the Construction

The most useful way to understand the construction is to follow an instance x. Each selected base hypothesis is applied separately to x. The results are collected into the vector ψ(x) = (h1(x), ..., hT(x)) in R^T. The final hypothesis does not apply its halfspace directly to the original x. It applies that halfspace to ψ(x), the representation created by the base hypotheses.

applyapplyapplyh1(x)h2(x)hT(x)input vectorproducexinput instanceh1h1(x)ψ(x)(h1(x), ..., hT(x))w-halfspacehomogeneousfinal hypothesismember of L(B,T)h2h2(x)hThT(x)
How does an input x move through the T base hypotheses to become the vector used by the final hypothesis?

Tracing a Symbolic Input

Let T be 3 and let h1, h2, and h3 be selected from B. Trace an input x through the composed construction.

Apply h1: The first selected base hypothesis produces h1(x).

Apply h2: The second selected base hypothesis produces h2(x).

Apply h3: The third selected base hypothesis produces h3(x).

Collect the outputs: The three outputs form ψ(x) = (h1(x), h2(x), h3(x)).

Apply the final halfspace: The homogeneous halfspace defined by w receives ψ(x), not the original x directly.

The input is transformed into a three-coordinate prediction vector before the final hypothesis acts.

Why AdaBoost Fits L(B,T)

AdaBoost repeatedly obtains weak hypotheses from the base learning process. When the hypotheses selected across the boosting rounds are named h1 through hT, AdaBoost has the same composition structure used to define L(B,T): the selected hypotheses produce a prediction vector, and a halfspace over that vector produces the final hypothesis. Therefore, the AdaBoost output belongs to L(B,T).

selectselectselectcontributescontributescontributescontrolsprediction sourceprediction sourceprediction sourceproducesBase class Bweak-hypothesis sourceRound 1h1wvector in R^Thalfspace compositionover base predictionsAdaBoost outputmember of L(B,T)Round 2h2Round ThT
What happens across the T boosting rounds so that the final AdaBoost hypothesis has the form required for membership in L(B,T)?

Mistakes About the Construction

  • Treating B as the final hypothesis class.

    B supplies the individual hypotheses, while L(B,T) describes the larger family formed by combining T of them with w.

    Fix: Use B for the building blocks and L(B,T) for the composed hypotheses.

  • Applying the final halfspace directly to x.

    The base hypotheses first transform x into ψ(x) = (h1(x), ..., hT(x)).

    Fix: State that the halfspace acts on ψ(x), the vector of base-hypothesis outputs.

  • Leaving out either the selected hypotheses or w.

    The construction is parameterized by both the T selected members of B and a vector w in R^T.

    Fix: Record both parts of the parameterization.

Check the Construction

MEDIUM

Suppose an algorithm selects T hypotheses h1 through hT from B and uses a vector w in R^T. Explain, in order, what happens to an input x and why the resulting final hypothesis is a member of L(B,T).

Hints
  • First identify the output of each selected base hypothesis on x.
  • Then write the prediction vector using those outputs.
  • Finally identify the operation controlled by w and compare the resulting structure with the definition of L(B,T).
  1. B is the restricted collection from which individual weak hypotheses are selected. L(B,T) is the larger class built from T selected hypotheses in B and a vector w in R^T. An input x is passed through each selected hypothesis to create ψ(x) = (h1(x), ..., hT(x)). A homogeneous halfspace defined by w then acts on ψ(x). AdaBoost has this same composition structure, so its output belongs to L(B,T).

Key Takeaways

  • B supplies the individual hypotheses used as building blocks.
  • A member of L(B,T) is determined by T selected hypotheses from B and a vector w in R^T.
  • The input x is transformed into ψ(x) = (h1(x), ..., hT(x)) before the final halfspace is applied.
  • AdaBoost belongs to L(B,T) because its output has this same composition structure.
  • The final composed hypothesis is not itself required to be an individual member of B.