AdaBoost
Model selection is the task of choosing an algorithm for a particular problem and setting its parameters.
The Decision Before Training
Suppose a practical problem could be approached with several algorithms, and each algorithm could be configured in several ways. The central challenge is not only writing down an algorithm. You must decide which algorithm fits the problem and which parameter settings fit it. This decision-making task is called model selection.
Model selection is the task of choosing an algorithm for a particular problem and setting its parameters. A selected model is therefore defined by both an algorithm choice and parameter choices.
Rounds of Adaptive Attention
AdaBoost, short for Adaptive Boosting, begins with access to a weak learner rather than a single powerful classifier. Across consecutive rounds, the weak learner receives a distribution over the training examples and produces a hypothesis from that distribution and the sample. AdaBoost then evaluates that hypothesis. Examples it misclassifies receive more attention in the following round, encouraging the next hypothesis to focus on cases that earlier hypotheses handled poorly.
What do you think happens?
A weak hypothesis handles some training examples poorly. What should AdaBoost do with those examples in the following round?
Reveal answer
Answer: Give them more attention
AdaBoost adapts by changing the distribution over training examples. Misclassified examples receive more attention in the following round, encouraging later learning to focus on them.
The Weak Learner's Input
The weak learner is the learning component AdaBoost repeatedly uses. At a given round, it works with the current distribution over training examples and the training sample, then produces a hypothesis. The distribution determines how much attention different examples receive at that point in the process. After the hypothesis is evaluated, its mistakes help determine the distribution used in the next round.
Tracing Two Rounds
Follow the role of one training example that the first weak hypothesis misclassifies.
Round 1 input: The weak learner receives the training sample together with the current distribution over training examples.
Round 1 output: The weak learner produces a hypothesis, and AdaBoost evaluates that hypothesis's error.
Attention update: Because the example was misclassified, it receives more attention in the following round.
Round 2 input: The weak learner now receives the changed distribution, so the next hypothesis is encouraged to focus more on examples handled poorly earlier.
AdaBoost is adaptive because the result of one round changes the distribution used in the next round.
Error Becomes Influence
AdaBoost does not simply keep every weak hypothesis with equal influence, nor does it automatically discard a hypothesis after measuring it. A hypothesis's error determines its weight in the final combination. Thus, each hypothesis has two consequences: its error affects how much influence it receives, and its mistakes affect which training examples receive more attention later.
The important distinction is between error and attention. The error of a hypothesis determines that hypothesis's weight in the final combination. Separately, the examples that the hypothesis misclassifies receive more attention in the following round. AdaBoost therefore evaluates each weak hypothesis individually while also using its mistakes to guide later learning.
Building the Strong Classifier
After successive rounds, AdaBoost has a collection of weak hypotheses and a weight for each one. It combines those weighted hypotheses to form the strong classifier. The final classifier is therefore not just the last weak hypothesis. It uses the collection of hypotheses, with each hypothesis contributing according to the influence assigned from its measured error.
Why the Collection Matters
Compare the role of one weak hypothesis with the role of the final strong classifier.
One hypothesis: A weak hypothesis is produced from the current distribution and sample. Its measured error determines its weight.
Several hypotheses: Across rounds, AdaBoost collects multiple weak hypotheses. Each one has an influence determined by its error.
Final combination: AdaBoost combines the weighted hypotheses rather than relying only on the most recently produced hypothesis.
The strong classifier is the weighted combination of the weak hypotheses gathered across rounds.
Choosing the Number of Rounds
AdaBoost's parameter T controls the number of rounds used in the process. Changing T changes how the model is configured, so selecting T belongs to model selection. The source identifies T as controlling the bias-complexity tradeoff: parameter setting requires considering the balance between the model's bias and its complexity.
Mistakes in Reasoning
Defining model selection as only choosing an algorithm.
Model selection also includes setting the algorithm's parameters.
Fix:
Include both decisions: choose an algorithm and choose its parameter settings.Treating T as a detail outside model selection.
T controls the bias-complexity tradeoff, so choosing T is part of selecting the model.
Fix:
Ask how T should be set for the particular practical problem.Assuming every weak hypothesis has equal influence.
A hypothesis's error determines its weight in the final combination.
Fix:
Track the measured error and the resulting hypothesis weight.Assuming AdaBoost focuses on the same examples in every round.
Misclassified examples receive more attention in the following round.
Fix:
Trace how each round's mistakes change the distribution used by the next round.Calling the last weak hypothesis the strong classifier.
AdaBoost forms the strong classifier by combining the weighted weak hypotheses.
Fix:
Describe the final classifier as a weighted combination of the collected hypotheses.
Check Your Understanding
Explain AdaBoost's process in four connected steps: first identify what the weak learner receives, then state what it produces, then explain how hypothesis error affects the hypothesis's weight, and finally describe how mistakes change the next example distribution and contribute to the strong classifier.
Hints
- Mention the training sample and the current distribution over training examples.
- Separate the hypothesis's weight from the increased attention given to misclassified examples.
- End with the weighted combination of weak hypotheses.
In your own words, explain why choosing AdaBoost's parameter T is a model-selection decision rather than a post-selection implementation detail.
Hints
- Start with the two parts of model selection.
- Connect T to the bias-complexity tradeoff.
- Mention that no universally best setting is given.
Key Takeaways
- Model selection includes choosing an algorithm and setting its parameters.
- AdaBoost uses a weak learner repeatedly across rounds and changes the distribution over training examples.
- Misclassified examples receive more attention in the following round.
- A hypothesis's measured error determines its weight in the final combination.
- AdaBoost forms a strong classifier by combining weighted weak hypotheses.
- The parameter T controls the bias-complexity tradeoff, so choosing T is part of model selection.
Key Takeaways
- Model selection is both algorithm choice and parameter setting.
- AdaBoost adapts from round to round by changing the distribution over training examples.
- Hypothesis error determines hypothesis influence, while misclassified examples receive more attention later.
- The final strong classifier combines the weighted weak hypotheses.
- AdaBoost's T belongs to model selection because it controls the bias-complexity tradeoff.