Statistical Error Bounds
The learning protocol separates hypothesis construction from performance estimation.
Separating Construction from Evaluation
A statistical error bound begins with a separation of responsibilities. One sequence of examples is used to construct a hypothesis, and another fresh sequence is used afterward to estimate how that hypothesis performs. This protocol prevents the examples used for construction from being confused with the examples used for performance estimation.
The Three Protocol Stages
The protocol can be read as three ordered stages. First, identify T, the sequence of k examples available during construction. Second, use T to construct the hypothesis h_T. The subscript T records that this hypothesis was produced from T. Third, use V, a separate sequence of m-k fresh examples, to calculate the error of h_T and then apply Bernstein's inequality to obtain a result related to that error.
- Use the k examples in T to construct h_T.
- Keep V separate from T; V contains m-k fresh examples.
- Evaluate h_T on V and use Bernstein's inequality to derive a bound related to that error.
Training and Fresh Sequences
T and V have different roles. T contains k examples and is available when h_T is constructed. V contains m-k examples and is fresh and independent of T. Because V is separate from the sequence that produced h_T, it can be used for performance estimation without changing the construction stage.
| Object | Size | Role |
|---|---|---|
| T | k examples | Used to construct h_T |
| h_T | One hypothesis | Produced from T |
| V | m-k examples | Used to calculate the error of h_T |
Why Evaluation Uses V
The protocol evaluates h_T on V because h_T was constructed from T. Reusing T would mix construction and performance estimation. In contrast, V is fresh and independent of T, so its examples were not part of the process that produced h_T. That independence supports the application of Bernstein's inequality.
When reading this protocol, ask two questions in order: Which sequence produced the hypothesis? Which separate sequence is used to estimate its error? The answers should be T and V, respectively.
A Protocol Walkthrough
Following the Named Objects
Suppose a protocol divides its examples into a sequence T of k examples and a fresh sequence V of m-k examples. Identify what happens in each stage.
Stage 1: construction data: T is the sequence available for construction. It contains k examples.
Stage 2: hypothesis construction: The learning procedure uses T to produce h_T. The subscript indicates the sequence used to construct the hypothesis.
Stage 3: fresh evaluation: The separate sequence V contains m-k fresh examples. The error of h_T is calculated on V rather than on T.
Bound: Because V and T are independent, Bernstein's inequality can be applied to obtain a result related to the error of h_T on V.
The protocol moves from T, to h_T, to the error of h_T on V, with Bernstein's inequality providing the error-related bound.
Changing the Construction Sequence
The notation h_T ties the hypothesis to the sequence that produced it. If the construction sequence changes, the constructed hypothesis can change as well. The evaluation sequence V remains conceptually separate: it is the fresh sequence used afterward to calculate the error of the hypothesis produced from the chosen construction sequence.
Common Mistakes
Treating T and V as interchangeable.
T contains the k examples used for construction, while V contains the m-k fresh examples used for evaluation.
Fix:
Remember the role assignment: T constructs h_T, and V is used to calculate the error of h_T.Evaluating h_T on T as the protocol's final performance estimate.
The protocol separates hypothesis construction from performance estimation.
Fix:
Evaluate h_T on V, the fresh sequence independent of T.Forgetting what the subscript in h_T records.
The notation identifies the construction sequence associated with the hypothesis.
Fix:
Read h_T as the hypothesis constructed from T.Applying Bernstein's inequality without preserving the separation.
The source identifies the independence of T and V as supporting the application of Bernstein's inequality.
Fix:
First identify T and V as separate, independent sequences; then relate Bernstein's inequality to the error of h_T on V.
Check Your Understanding
A protocol uses T to construct h_T and then uses a fresh sequence V to calculate the error of h_T. Explain the role of each object and state why Bernstein's inequality is relevant.
Hints
- State how many examples T and V contain.
- Explain what the subscript T indicates.
- Use the independence of T and V when explaining Bernstein's inequality.
What do you think happens?
Which sequence should be used to calculate the error of h_T in this protocol?
Reveal answer
Answer: V, because it is a fresh sequence of m-k examples independent of T.
T is used to construct h_T. V is reserved for performance estimation, and its independence from T supports the application of Bernstein's inequality.
Protocol Summary
- T contains k examples and is used to construct h_T.
- V contains m-k fresh examples and is independent of T.
- The error of h_T is calculated on V rather than on T.
- The separation between T and V supports applying Bernstein's inequality to obtain a result related to the error of h_T on V.
Key Takeaways
- The protocol has three stages: use T for construction, produce h_T, and evaluate h_T on fresh V.
- T contains k examples, while V contains m-k fresh examples.
- Evaluating on V keeps performance estimation separate from hypothesis construction.
- The independence of T and V supports the use of Bernstein's inequality for a bound related to the error of h_T on V.