Concepts / AdaBoost

AdaBoost

Model selection is the task of choosing an algorithm for a particular problem and setting its parameters.

  • Programming

The Decision Before Training

Suppose a practical problem could be approached with several algorithms, and each algorithm could be configured in several ways. The central challenge is not only writing down an algorithm. You must decide which algorithm fits the problem and which parameter settings fit it. This decision-making task is called model selection.

choosesetexampleincludesPractical problemAlgorithm choiceAdaBoostParameter choicesTnumber of rounds
What two choices define the selected model, and where does AdaBoost's T belong?

Model selection is the task of choosing an algorithm for a particular problem and setting its parameters. A selected model is therefore defined by both an algorithm choice and parameter choices.

Rounds of Adaptive Attention

AdaBoost, short for Adaptive Boosting, begins with access to a weak learner rather than a single powerful classifier. Across consecutive rounds, the weak learner receives a distribution over the training examples and produces a hypothesis from that distribution and the sample. AdaBoost then evaluates that hypothesis. Examples it misclassifies receive more attention in the following round, encouraging the next hypothesis to focus on cases that earlier hypotheses handled poorly.

guidesis evaluated bychanges attentionguidesExample distributionround 1Weak hypothesisround 1Hypothesis errorround 1Example distributionround 2Weak hypothesisround 2
What happens to attention on training examples after each weak learner makes mistakes?

What do you think happens?

A weak hypothesis handles some training examples poorly. What should AdaBoost do with those examples in the following round?

  • Give them more attention
  • Give them less attention
  • Discard the entire training set
Reveal answer

Answer: Give them more attention

AdaBoost adapts by changing the distribution over training examples. Misclassified examples receive more attention in the following round, encouraging later learning to focus on them.

The Weak Learner's Input

The weak learner is the learning component AdaBoost repeatedly uses. At a given round, it works with the current distribution over training examples and the training sample, then produces a hypothesis. The distribution determines how much attention different examples receive at that point in the process. After the hypothesis is evaluated, its mistakes help determine the distribution used in the next round.

withguidesproducesis evaluated byupdates attentionTraining sampleWeak learnerHypothesisHypothesis errorExample distributionnext roundExampledistributioncurrent round
How does the current distribution flow into the weak learner, and how does the learner's output affect the next distribution?

Tracing Two Rounds

Follow the role of one training example that the first weak hypothesis misclassifies.

Round 1 input: The weak learner receives the training sample together with the current distribution over training examples.

Round 1 output: The weak learner produces a hypothesis, and AdaBoost evaluates that hypothesis's error.

Attention update: Because the example was misclassified, it receives more attention in the following round.

Round 2 input: The weak learner now receives the changed distribution, so the next hypothesis is encouraged to focus more on examples handled poorly earlier.

AdaBoost is adaptive because the result of one round changes the distribution used in the next round.

Error Becomes Influence

AdaBoost does not simply keep every weak hypothesis with equal influence, nor does it automatically discard a hypothesis after measuring it. A hypothesis's error determines its weight in the final combination. Thus, each hypothesis has two consequences: its error affects how much influence it receives, and its mistakes affect which training examples receive more attention later.

determinesmistakes affectHypothesis errormeasured on weightedtraining setHypothesis weightin final combinationWeak hypothesisExample attentionnext round
How does a weak hypothesis's error on the weighted training set determine its influence in the final classifier?

The important distinction is between error and attention. The error of a hypothesis determines that hypothesis's weight in the final combination. Separately, the examples that the hypothesis misclassifies receive more attention in the following round. AdaBoost therefore evaluates each weak hypothesis individually while also using its mistakes to guide later learning.

Building the Strong Classifier

After successive rounds, AdaBoost has a collection of weak hypotheses and a weight for each one. It combines those weighted hypotheses to form the strong classifier. The final classifier is therefore not just the last weak hypothesis. It uses the collection of hypotheses, with each hypothesis contributing according to the influence assigned from its measured error.

contributescontributescontributesproducesWeak hypothesis 1weight 1Weighted combinationStrong classifierfinal predictionWeak hypothesis 2weight 2Weak hypothesis 3weight 3
How are individual weak hypotheses and their weights combined to produce the final prediction?

Why the Collection Matters

Compare the role of one weak hypothesis with the role of the final strong classifier.

One hypothesis: A weak hypothesis is produced from the current distribution and sample. Its measured error determines its weight.

Several hypotheses: Across rounds, AdaBoost collects multiple weak hypotheses. Each one has an influence determined by its error.

Final combination: AdaBoost combines the weighted hypotheses rather than relying only on the most recently produced hypothesis.

The strong classifier is the weighted combination of the weak hypotheses gathered across rounds.

Choosing the Number of Rounds

AdaBoost's parameter T controls the number of rounds used in the process. Changing T changes how the model is configured, so selecting T belongs to model selection. The source identifies T as controlling the bias-complexity tradeoff: parameter setting requires considering the balance between the model's bias and its complexity.

model decision includescontrolsAdaBoostalgorithm choiceBias-complexitytradeoffTparameter choice
How does changing AdaBoost's number of rounds T relate to model complexity and the bias-complexity tradeoff?

Mistakes in Reasoning

  • Defining model selection as only choosing an algorithm.

    Model selection also includes setting the algorithm's parameters.

    Fix: Include both decisions: choose an algorithm and choose its parameter settings.

  • Treating T as a detail outside model selection.

    T controls the bias-complexity tradeoff, so choosing T is part of selecting the model.

    Fix: Ask how T should be set for the particular practical problem.

  • Assuming every weak hypothesis has equal influence.

    A hypothesis's error determines its weight in the final combination.

    Fix: Track the measured error and the resulting hypothesis weight.

  • Assuming AdaBoost focuses on the same examples in every round.

    Misclassified examples receive more attention in the following round.

    Fix: Trace how each round's mistakes change the distribution used by the next round.

  • Calling the last weak hypothesis the strong classifier.

    AdaBoost forms the strong classifier by combining the weighted weak hypotheses.

    Fix: Describe the final classifier as a weighted combination of the collected hypotheses.

Check Your Understanding

MEDIUM

Explain AdaBoost's process in four connected steps: first identify what the weak learner receives, then state what it produces, then explain how hypothesis error affects the hypothesis's weight, and finally describe how mistakes change the next example distribution and contribute to the strong classifier.

Hints
  • Mention the training sample and the current distribution over training examples.
  • Separate the hypothesis's weight from the increased attention given to misclassified examples.
  • End with the weighted combination of weak hypotheses.
EASY

In your own words, explain why choosing AdaBoost's parameter T is a model-selection decision rather than a post-selection implementation detail.

Hints
  • Start with the two parts of model selection.
  • Connect T to the bias-complexity tradeoff.
  • Mention that no universally best setting is given.

Key Takeaways

  1. Model selection includes choosing an algorithm and setting its parameters.
  2. AdaBoost uses a weak learner repeatedly across rounds and changes the distribution over training examples.
  3. Misclassified examples receive more attention in the following round.
  4. A hypothesis's measured error determines its weight in the final combination.
  5. AdaBoost forms a strong classifier by combining weighted weak hypotheses.
  6. The parameter T controls the bias-complexity tradeoff, so choosing T is part of model selection.

Key Takeaways

  • Model selection is both algorithm choice and parameter setting.
  • AdaBoost adapts from round to round by changing the distribution over training examples.
  • Hypothesis error determines hypothesis influence, while misclassified examples receive more attention later.
  • The final strong classifier combines the weighted weak hypotheses.
  • AdaBoost's T belongs to model selection because it controls the bias-complexity tradeoff.