Boosting Weak Learners
The task is binary face recognition: classify an image as a human face or not.
From Image to Decision
The source material frames face recognition as a binary classification task. Given an image, the learned classifier must decide between two classes: human face and not a human face. The important idea is not only the final decision, but also the sequence of representations that leads to it: the image is processed by a rectangle-based function g, that function produces a scalar, and a function f uses that scalar within the base hypothesis h.
A weak learner in this setting begins with one image, extracts one scalar associated with a selected rectangle, and uses that scalar to form a classification hypothesis.
The 24 × 24 Input
The input space consists of real-valued 24 × 24 image matrices. An individual input instance x is therefore one such matrix: it has 24 positions across and 24 positions down, giving 576 pixel positions in total. At this stage, x is still the complete image. It is not yet the scalar used by the base hypothesis and it is not yet a face-recognition output.
Identifying the Input Instance
A face-recognition example receives one real-valued 24 × 24 image matrix. What is the input x?
Step 1: Treat the complete 24 × 24 matrix as one input instance, called x.
Step 2: Recognize that x contains 576 pixel positions, because 24 positions in each of two dimensions produce 24 × 24 positions.
Step 3: Keep x distinct from any later scalar produced from a selected rectangle.
The input instance is the complete image matrix x, not a single pixel, not a rectangle by itself, and not the final binary classification.
Selecting a Rectangle
The function g is parameterized by an axis-aligned rectangle R. Axis-aligned means the rectangle is specified within the image's row-and-column orientation rather than being described as an arbitrary rotated region. Its parameters determine which rectangular region of the 24 × 24 image is selected. The source states that g maps the image to a scalar associated with that selected rectangle. It does not specify a particular operation such as a sum or an average, so the essential point is the selection and the resulting scalar, not an assumed calculation.
The rectangle is a parameter of g, not a replacement for the image. The image remains x; the rectangle specifies how g obtains its scalar from x.
For a 24 × 24 image, the source states that there are at most 24^4 axis-aligned rectangles. These rectangles form a finite collection of possible parameters for functions g in the base hypothesis class.
Following h(x)
The base hypothesis is written as h(x) = f(g(x)). Read this from the inside outward. First, g receives the full image x and produces one scalar associated with a selected axis-aligned rectangle. Next, f receives that scalar as part of the base hypothesis. In the source's terminology, f is the decision-stump component. The complete function h is the resulting hypothesis used for the face-recognition task.
Tracing One Base Hypothesis
Trace the roles of x, g, f, and h in h(x) = f(g(x)).
Input: x is the complete real-valued 24 × 24 image matrix.
Rectangle-based processing: g uses an axis-aligned rectangle parameter R and maps the image to one scalar associated with that rectangle.
Decision-stump stage: f uses the scalar produced by g as the input to the decision-stump component.
Complete hypothesis: h combines these stages and is the base hypothesis for the binary face-recognition task.
The expression separates the image representation, scalar-producing function, decision-stump component, and complete hypothesis.
Keeping the Stages Separate
| Stage | Role | Representation |
|---|---|---|
| Image input x | The complete item being classified | Real-valued 24 × 24 matrix |
| Function g | Uses a rectangle parameter and maps the image to a scalar | Rectangle-based scalar-producing function |
| Function f | The decision-stump component that uses the scalar | Function applied after g |
| Hypothesis h(x) | The complete base hypothesis for the task | Composition f(g(x)) |
Common Interpretation Errors
Treating x as a single pixel or as the selected rectangle.
The input is the complete real-valued 24 × 24 image matrix. The rectangle is a parameter used by g.
Fix:
Keep x as the full image and R as the axis-aligned rectangle that parameterizes g.Treating g(x) as the final classifier output.
The decomposition continues with f: h(x) = f(g(x)).
Fix:
Identify g(x) as the scalar passed to the decision-stump component f.Replacing the composition with an unspecified image operation.
The source states that g maps the image to a scalar associated with a selected rectangle, but it does not specify that particular calculation.
Fix:
Describe only the stated behavior: an axis-aligned rectangle parameterizes g, which maps the image to a scalar.Forgetting that the classification task has two classes.
The source frames the task as deciding whether the image is a human face or not a human face.
Fix:
Use the binary labels human face and not a human face when describing the task.
Check Your Understanding
Explain the path of one 24 × 24 image through the base hypothesis. In your explanation, name the input x, the axis-aligned rectangle R, the scalar g(x), the decision-stump component f, and the complete hypothesis h(x).
Hints
- Begin with the complete image matrix, not the rectangle.
- State what g produces before describing f.
- End by connecting the composition to the binary human-face versus not-a-human-face task.
What do you think happens?
Before checking the explanation, which expression correctly preserves the order of the base hypothesis stages?
Reveal answer
Answer: h(x) = f(g(x))
The image x is first passed to g, which produces a scalar. The function f then uses that scalar, giving the composition f(g(x)).
Key Takeaways
- The task is binary face recognition: classify an image as a human face or not a human face.
- Each input x is a real-valued 24 × 24 image matrix with 576 pixel positions.
- An axis-aligned rectangle R parameterizes g, which maps the complete image to a scalar.
- The base hypothesis is composed as h(x) = f(g(x)); f is the decision-stump component applied after g.
- For a 24 × 24 image, the source states that there are at most 24^4 axis-aligned rectangles available as parameters for functions g.
Key Takeaways
- Face recognition here is a binary decision between human face and not a human face.
- The input x is the full real-valued 24 × 24 image matrix.
- The function g uses an axis-aligned rectangle parameter to produce one scalar from x.
- The decision-stump component f acts after g, so the complete base hypothesis is h(x) = f(g(x)).
- The rectangle parameters create a finite collection of possible functions g, with at most 24^4 axis-aligned rectangles for a 24 × 24 image.