Concepts / Model Selection

Model Selection

Validation uses some training data as a validation set.

  • Programming

Why Validation Is Needed

A learning algorithm produces an output predictor, but producing that predictor is not the end of the task. We also want to know how well the predictor will perform in terms of its true risk. The available evidence may not provide an estimate that is as informative as we would like, so a separate validation procedure can provide additional information.

Validation uses some of the original training data as a validation set. The output predictor is evaluated on that set, and its observed success supplies information for estimating true risk. In this way, validation gives us evidence about the predictor beyond simply knowing that a learning algorithm was run.

part ofsome set asidelearning algorithmevaluated onOriginal trainingdataData used to fitOutput predictorValidation successValidation set
How does the original training data get divided, and which portion is used to fit the model versus evaluate the predictor?

Tracing a Validation Estimate

Comparing Two Candidate Predictors

Suppose a practical problem produces two candidate output predictors. How can validation help decide which candidate to prefer?

Produce candidates: Learning procedures produce two candidate output predictors for the same practical problem.

Use the same validation procedure: Evaluate both candidates using the same validation procedure and validation set.

Compare validation success: Compare the observed validation success of the candidates.

Estimate true risk: The validation results provide information for estimating the candidates' true risks.

Prefer the stronger result: The candidate with stronger validation success can be preferred because validation gives a more informative estimate than relying only on a broad, potentially pessimistic bound.

Validation turns observed performance on the validation set into evidence that can support a model choice.

producesevaluated onprovides evaluation datasupplies information forLearning algorithmOutput predictorObserved successTrue risk estimateValidation set
How does validation performance provide an estimate of how well a predictor will perform on unseen data?

Why Broad Bounds Can Be Limited

One way to reason about a hypothesis class is to use bounds on estimation error. Such bounds state that, for every hypothesis in the class, true risk is not very far from empirical risk. Their limitation is that they must cover all hypotheses and all possible data distributions. Because they provide this broad coverage, the bounds can be loose and pessimistic.

Validation changes the source of the estimate. Instead of relying only on a broad bound, the procedure evaluates the particular output predictor on the validation set. The observed result is then used as information for estimating that predictor's true risk. This more targeted evidence is why validation is useful when selecting among candidate models.

The Two Choices in Model Selection

Model selection is the task of choosing an algorithm for a particular problem and setting its parameters.

The definition contains two connected decisions. First, choose which learning algorithm should be used. Second, decide how that algorithm should be configured by choosing its parameter settings. Model selection is therefore not merely asking which algorithm to use. It also asks which settings fit the practical problem.

requires a decision aboutrequires a decision aboutcombines withconfiguresis evaluated throughsupports selection ofPractical problemAlgorithm choiceCandidate modelValidation evidenceSelected modelParameter choices
How do the choices of learning algorithm and model parameters combine to determine the final model?
DecisionQuestion it answersPlace in model selection
Algorithm choiceWhich learning algorithm should address the problem?One part of selecting a model
Parameter choiceHow should the chosen algorithm be configured?The other part of selecting a model

A selected model is defined by both an algorithm choice and parameter choices.

AdaBoost and Parameter T

AdaBoost illustrates why parameter setting belongs inside model selection. Its parameter T controls the bias-complexity tradeoff. Therefore, deciding how to set T is not a minor step that happens after model selection. It is part of selecting the model itself.

setsinvolvescan be evaluated throughChoose TAdaBoost modelparameterized by TBias-complexitytradeoffValidation result
How does considering different values of T place AdaBoost parameter setting inside the model selection problem?

The source does not identify one universally best value for T. The appropriate setting is tied to the practical problem at hand. This is precisely why T belongs in the selection question: different parameter choices represent different candidate configurations whose suitability can be considered using validation.

Balancing Bias and Complexity

Parameter selection can involve a bias-complexity tradeoff. In the source, AdaBoost's T is the concrete parameter associated with this tradeoff. The important practical lesson is not to treat parameter values as an afterthought. A parameter setting helps define the model, so choosing it is part of model selection.

must be consideredmust be consideredinformsParameter setting Aone candidate configurationBias-complexitytradeoffModel selectionParameter setting Banother candidateconfiguration
What consideration must be kept in view when a parameter setting changes the model's bias-complexity balance?

Common Reasoning Mistakes

  • Treating model selection as only the choice of an algorithm.

    Model selection includes both choosing an algorithm and setting its parameters.

    Fix: Ask two questions: which algorithm fits the problem, and how should that algorithm be configured?

  • Treating validation as the same thing as a broad estimation-error bound.

    Broad bounds can be loose and pessimistic because they must provide coverage so widely.

    Fix: Recognize that validation evaluates the particular output predictor on a validation set and supplies information for a more accurate true-risk estimate.

  • Treating AdaBoost's T as an implementation detail outside model selection.

    T controls the bias-complexity tradeoff, so deciding its value is part of choosing the model.

    Fix: Include T among the parameter choices considered during model selection.

  • Claiming that one parameter setting is universally best.

    The appropriate selection is tied to the practical problem at hand.

    Fix: Use validation evidence and the problem context when comparing candidate configurations.

Apply the Selection Logic

MEDIUM

A practical problem has two candidate predictors. You have set aside some original training data as a validation set. In your own words, explain what the validation results contribute to the decision, and identify the two kinds of choices that together define the selected model.

Hints
  • Start with what is evaluated on the validation set.
  • Connect the observed validation success to an estimate of true risk.
  • Name both the algorithm decision and the parameter decision.

What do you think happens?

A learner says, “I have finished model selection once I choose AdaBoost; choosing T is separate.” Is that statement correct?

  • Yes, because T is only an implementation detail.
  • No, because T is a parameter whose setting is part of selecting the model.
Reveal answer

Answer: No, because T is a parameter whose setting is part of selecting the model.

AdaBoost's T controls the bias-complexity tradeoff. Since model selection includes setting parameters, deciding T belongs to the model selection problem.

Key Takeaways

  1. Validation sets use some of the original training data to evaluate an output predictor.
  2. Validation success supplies information for estimating true risk and can be more informative than broad, loose, or pessimistic bounds.
  3. Model selection means choosing both a learning algorithm and its parameter settings.
  4. AdaBoost's parameter T belongs to model selection because it controls the bias-complexity tradeoff.
  5. Parameter choices should be considered in relation to the particular practical problem; no universally best setting is given.

Key Takeaways

  • Validation evaluates a predictor on some training data set aside as a validation set.
  • The validation result provides a more informative estimate of true risk than relying only on broad, potentially pessimistic bounds.
  • Model selection includes choosing an algorithm and setting its parameters.
  • AdaBoost's T is part of model selection because it controls the bias-complexity tradeoff.
  • The best selection depends on the particular practical problem.