Concepts / Boosting Weak Learners

Boosting Weak Learners

The task is binary face recognition: classify an image as a human face or not.

  • Programming

From Image to Decision

The source material frames face recognition as a binary classification task. Given an image, the learned classifier must decide between two classes: human face and not a human face. The important idea is not only the final decision, but also the sequence of representations that leads to it: the image is processed by a rectangle-based function g, that function produces a scalar, and a function f uses that scalar within the base hypothesis h.

A weak learner in this setting begins with one image, extracts one scalar associated with a selected rectangle, and uses that scalar to form a classification hypothesis.

The 24 × 24 Input

The input space consists of real-valued 24 × 24 image matrices. An individual input instance x is therefore one such matrix: it has 24 positions across and 24 positions down, giving 576 pixel positions in total. At this stage, x is still the complete image. It is not yet the scalar used by the base hypothesis and it is not yet a face-recognition output.

hashascombined with columnsforms positionsx24 × 24 image matrixRow positions24 positionsPixel positions576 total positionsColumn positions24 positions
How do the 576 pixel positions in a 24 × 24 image make up the input instance x?

Identifying the Input Instance

A face-recognition example receives one real-valued 24 × 24 image matrix. What is the input x?

Step 1: Treat the complete 24 × 24 matrix as one input instance, called x.

Step 2: Recognize that x contains 576 pixel positions, because 24 positions in each of two dimensions produce 24 × 24 positions.

Step 3: Keep x distinct from any later scalar produced from a selected rectangle.

The input instance is the complete image matrix x, not a single pixel, not a rectangle by itself, and not the final binary classification.

Selecting a Rectangle

The function g is parameterized by an axis-aligned rectangle R. Axis-aligned means the rectangle is specified within the image's row-and-column orientation rather than being described as an arbitrary rotated region. Its parameters determine which rectangular region of the 24 × 24 image is selected. The source states that g maps the image to a scalar associated with that selected rectangle. It does not specify a particular operation such as a sum or an average, so the essential point is the selection and the resulting scalar, not an assumed calculation.

selects regionparameterizesImage x24 × 24 matrixRectangle Rposition and dimensionsg(x)one scalar
Which rectangular region does g select, and how do the rectangle's position and dimensions determine the scalar-producing function?

The rectangle is a parameter of g, not a replacement for the image. The image remains x; the rectangle specifies how g obtains its scalar from x.

For a 24 × 24 image, the source states that there are at most 24^4 axis-aligned rectangles. These rectangles form a finite collection of possible parameters for functions g in the base hypothesis class.

Following h(x)

The base hypothesis is written as h(x) = f(g(x)). Read this from the inside outward. First, g receives the full image x and produces one scalar associated with a selected axis-aligned rectangle. Next, f receives that scalar as part of the base hypothesis. In the source's terminology, f is the decision-stump component. The complete function h is the resulting hypothesis used for the face-recognition task.

receivesproducessuppliesformsximage matrixgrectangle-based functiong(x)one scalarfdecision stumph(x)hypothesis output
How does the image x pass through g and then f to produce the hypothesis output h(x)?

Tracing One Base Hypothesis

Trace the roles of x, g, f, and h in h(x) = f(g(x)).

Input: x is the complete real-valued 24 × 24 image matrix.

Rectangle-based processing: g uses an axis-aligned rectangle parameter R and maps the image to one scalar associated with that rectangle.

Decision-stump stage: f uses the scalar produced by g as the input to the decision-stump component.

Complete hypothesis: h combines these stages and is the base hypothesis for the binary face-recognition task.

The expression separates the image representation, scalar-producing function, decision-stump component, and complete hypothesis.

Keeping the Stages Separate

g mapsf receivescontributes toImage x24 × 24 matrixScalar g(x)rectangle-associated valueDecision stump fuses the scalarHypothesis h(x)face-recognition output
What changes at each stage from the image input to the final classification hypothesis?
StageRoleRepresentation
Image input xThe complete item being classifiedReal-valued 24 × 24 matrix
Function gUses a rectangle parameter and maps the image to a scalarRectangle-based scalar-producing function
Function fThe decision-stump component that uses the scalarFunction applied after g
Hypothesis h(x)The complete base hypothesis for the taskComposition f(g(x))

Common Interpretation Errors

  • Treating x as a single pixel or as the selected rectangle.

    The input is the complete real-valued 24 × 24 image matrix. The rectangle is a parameter used by g.

    Fix: Keep x as the full image and R as the axis-aligned rectangle that parameterizes g.

  • Treating g(x) as the final classifier output.

    The decomposition continues with f: h(x) = f(g(x)).

    Fix: Identify g(x) as the scalar passed to the decision-stump component f.

  • Replacing the composition with an unspecified image operation.

    The source states that g maps the image to a scalar associated with a selected rectangle, but it does not specify that particular calculation.

    Fix: Describe only the stated behavior: an axis-aligned rectangle parameterizes g, which maps the image to a scalar.

  • Forgetting that the classification task has two classes.

    The source frames the task as deciding whether the image is a human face or not a human face.

    Fix: Use the binary labels human face and not a human face when describing the task.

Check Your Understanding

MEDIUM

Explain the path of one 24 × 24 image through the base hypothesis. In your explanation, name the input x, the axis-aligned rectangle R, the scalar g(x), the decision-stump component f, and the complete hypothesis h(x).

Hints
  • Begin with the complete image matrix, not the rectangle.
  • State what g produces before describing f.
  • End by connecting the composition to the binary human-face versus not-a-human-face task.

What do you think happens?

Before checking the explanation, which expression correctly preserves the order of the base hypothesis stages?

  • h(x) = g(f(x))
  • h(x) = f(g(x))
  • h(x) = x(f(g))
Reveal answer

Answer: h(x) = f(g(x))

The image x is first passed to g, which produces a scalar. The function f then uses that scalar, giving the composition f(g(x)).

Key Takeaways

  1. The task is binary face recognition: classify an image as a human face or not a human face.
  2. Each input x is a real-valued 24 × 24 image matrix with 576 pixel positions.
  3. An axis-aligned rectangle R parameterizes g, which maps the complete image to a scalar.
  4. The base hypothesis is composed as h(x) = f(g(x)); f is the decision-stump component applied after g.
  5. For a 24 × 24 image, the source states that there are at most 24^4 axis-aligned rectangles available as parameters for functions g.

Key Takeaways

  • Face recognition here is a binary decision between human face and not a human face.
  • The input x is the full real-valued 24 × 24 image matrix.
  • The function g uses an axis-aligned rectangle parameter to produce one scalar from x.
  • The decision-stump component f acts after g, so the complete base hypothesis is h(x) = f(g(x)).
  • The rectangle parameters create a finite collection of possible functions g, with at most 24^4 axis-aligned rectangles for a 24 × 24 image.