Validation and Model Evaluation
The usefulness of a learnability definition depends on the goal for which it is being used.
Start with the Analytical Goal
A learnability definition is useful only in relation to the goal of the analysis. A formal name by itself does not tell you whether the definition gives the conclusion you need. In model evaluation, an important goal is to say something meaningful about the risk of the predictor produced by a learning algorithm.
The first question should therefore be: what must the analysis establish? If the goal is to constrain true risk from above, the analysis needs a learnability notion that supplies an upper-bound guarantee. If the immediate goal is to assess the risk of an already produced predictor, a validation set provides an estimation route. These are related but different analytical jobs.
From Validation Data to a Risk Estimate
Imagine that a learning algorithm has already produced an output predictor. The next question is how well that predictor performs beyond the particular data used during learning. A validation set can be used to estimate the risk of this output predictor.
The validation route gives evidence about the predictor's risk. It is described as an estimate because it assesses the output predictor using validation information. That role is different from the learnability guarantee supplied by PAC learning or nonuniform learning.
Three Ways to Read a Risk Claim
PAC learning, nonuniform learning, and consistency guarantees do not support the same conclusion about the output predictor. For the specific goal of bounding the risk of the learned hypothesis, PAC learning and nonuniform learning provide an upper bound on true risk. Consistency guarantees do not provide such an upper bound.
| Learnability notion or method | What it provides | Upper bound on true risk? |
|---|---|---|
| PAC learning | An upper bound on the true risk of the learned hypothesis | Yes |
| Nonuniform learning | An upper bound on the true risk of the learned hypothesis | Yes |
| Consistency guarantee | A consistency guarantee, but not an upper bound on true risk | No |
| Validation set | An estimate of the risk of the output predictor | Not the learnability upper-bound guarantee described above |
Reading the Output Predictor
Choosing the Right Conclusion
An analysis has produced an output predictor. The analyst wants to assess its risk immediately, and also wants to know which learnability notions can support a statement that true risk is no greater than a stated limit.
Separate the two goals: Assessing the output predictor's risk and bounding true risk are different analytical jobs.
Choose an assessment route: Use a validation set to estimate the risk of the output predictor.
Choose a bounding guarantee: PAC learning and nonuniform learning provide an upper bound on the true risk of the learned hypothesis.
Reject an unsupported conclusion: A consistency guarantee does not provide an upper bound on true risk.
Validation supplies an estimate for assessing the output predictor. PAC learning and nonuniform learning supply the upper-bound conclusion. Consistency alone does not supply that upper bound.
When you read a model-evaluation statement, identify its verb. Estimate indicates assessment of risk. Provides an upper bound indicates a stronger type of conclusion about what the true risk cannot exceed.
Common Interpretation Mistakes
Treating every learnability definition as equally useful for every goal
The usefulness of a learnability definition depends on the goal for which it is being used.
Fix:
Name the desired conclusion before choosing the definition.Calling a validation estimate an upper bound on true risk
A validation set estimates the risk of the output predictor; the source distinguishes that estimate from the upper-bound guarantee supplied by PAC or nonuniform learning.
Fix:
Describe validation as an estimation route unless a separate upper-bound guarantee is available.Assuming a consistency guarantee supplies an upper bound on true risk
Consistency guarantees do not provide an upper bound on true risk.
Fix:
Use PAC learning or nonuniform learning when the required conclusion is an upper bound on true risk.Focusing only on whether a hypothesis can be learned
The practical question is also whether the chosen definition helps control or assess the output predictor's true risk.
Fix:
Connect the learnability notion to the risk conclusion needed by the analysis.
Practice the Decision
For each need below, identify whether it calls for a validation estimate or an upper-bound guarantee. Then identify which learnability notions supply the upper-bound guarantee described in the source.
Hints
- Look for the difference between assessing an output predictor and constraining true risk from above.
- PAC learning and nonuniform learning provide an upper bound on true risk.
- Consistency guarantees do not provide an upper bound on true risk.
What do you think happens?
An analyst has an output predictor and wants immediate evidence about its risk. Which route matches that goal?
Reveal answer
Answer: Use a validation set to estimate the predictor's risk
The source identifies validation as an estimation route for the risk of an output predictor. It separately identifies PAC learning and nonuniform learning as notions that provide an upper bound on true risk.
Summary
- The usefulness of a learnability definition depends on the goal of the analysis.
- A validation set can estimate the risk of an output predictor.
- An estimate of risk is different from an upper bound on true risk.
- PAC learning and nonuniform learning provide an upper bound on true risk.
- Consistency guarantees do not provide an upper bound on true risk.
Key Takeaways
- Choose a learnability definition based on the conclusion the analysis must establish.
- Use validation to estimate the risk of an already produced output predictor.
- Do not confuse validation evidence with an upper-bound guarantee on true risk.
- PAC learning and nonuniform learning provide upper bounds on true risk.
- Consistency guarantees do not provide upper bounds on true risk.