Empirical Loss
The PAC-Bayes bound can be converted into a learning rule.
From Guarantee to Decision
A bound is often presented as a statement about how well a model may perform. The PAC-Bayes bound does more than provide such a statement: it can be converted into a learning rule. The rule begins with a prior distribution P, evaluates possible posterior distributions Q using the function associated with the bound, and returns a posterior that minimizes that function.
Tracing P and Q
The prior P is the distribution supplied to the learning rule at the beginning. The posterior Q is the distribution the rule must choose and return. The rule considers possible posterior distributions in relation to the given prior, evaluates the function associated with the PAC-Bayes bound, and selects a posterior that minimizes that function.
A qualitative choice among posteriors
Suppose a learning rule starts with a prior P and considers three possible posterior distributions: Q1, Q2, and Q3. Each posterior is evaluated with the function associated with the PAC-Bayes bound.
Start with P: P is fixed as the prior supplied to the rule.
Consider Q1, Q2, and Q3: These are candidate posterior distributions that the rule may evaluate.
Evaluate the candidates: The rule evaluates the specified function for each candidate posterior.
Return the minimum: The posterior with the smallest value of the specified function is returned.
The learning rule returns a posterior Q selected by minimization; the rule is not merely reporting the PAC-Bayes bound.
What do you think happens?
If two candidate posteriors have different values for the function associated with the PAC-Bayes bound, which one does the learning rule return?
Reveal answer
Answer: The posterior with the smallest function value
The PAC-Bayes bound is converted into a rule that evaluates candidate posteriors and returns one that minimizes the specified function.
The Regularized Objective
Regularized risk minimization combines two considerations. The first is empirical loss, which represents the model's loss on the observed training examples. The second is the Kullback-Leibler distance between Q and P, which supplies the distribution-comparison component of the objective. Together, these quantities define the trade-off considered when selecting a posterior.
| Quantity | What it measures | Role in the objective |
|---|---|---|
| Empirical loss | Loss on the observed training examples | Measures the data-fitting component |
| KL distance between Q and P | The distribution-comparison relationship between the posterior and prior | Provides the distribution-comparison component |
Regularized risk minimization considers both empirical loss and the KL distance between Q and P.
Two Quantities, One Choice
The learning rule must make one posterior choice while considering both parts of the regularized objective. A posterior with a favorable empirical loss is not automatically the selected posterior. Its relationship to the prior also matters through the KL-distance term. Conversely, a posterior's relationship to the prior is not the only consideration, because empirical loss is also part of the objective.
Why the combined function matters
Imagine two candidate posteriors, Q1 and Q2. Q1 is attractive because of its empirical loss. Q2 is attractive because of its relationship to the prior. The learning rule evaluates both candidates using the combined function associated with regularized risk minimization.
Inspect empirical loss: The rule considers how the candidates perform on the observed training examples.
Inspect the KL distance: The rule also considers the distribution-comparison relationship between each candidate posterior and the prior P.
Combine the considerations: Regularized risk minimization considers empirical loss together with the KL-distance component rather than using only one of them.
Select Q: The returned posterior is the candidate that minimizes the combined function.
The selected posterior is determined by the combined objective, not by empirical loss or KL distance considered in isolation.
Common Misreadings
Treating the PAC-Bayes bound only as a performance statement
The source describes a learning rule derived from the bound: given P, evaluate possible Q distributions and return one that minimizes a specified function.
Fix:
Read the bound as both a guarantee and a basis for an optimization procedure.Confusing the prior P with the returned posterior Q
P is given to the rule, while Q is the posterior the rule returns.
Fix:
Track the direction of the procedure: start with P, consider candidate Q values, then return a minimizing Q.Minimizing empirical loss alone
Regularized risk minimization combines empirical loss with the KL distance between Q and P.
Fix:
Evaluate the combined objective containing both quantities.Treating KL distance as the same thing as empirical loss
Empirical loss concerns the observed examples, while KL distance is the distribution-comparison component involving Q and P.
Fix:
Keep the data-fitting role of empirical loss separate from the distribution-comparison role of KL distance.
When explaining the learning rule, name all three stages explicitly: identify the prior P, evaluate candidate posteriors Q with the specified function, and return a minimizing Q. When explaining regularized risk minimization, name both minimized quantities: empirical loss and KL distance between Q and P.
Check Your Understanding
Explain in your own words how the PAC-Bayes bound becomes a learning rule. Your explanation should identify the input distribution, the candidate distributions, the function being evaluated, and the selection criterion.
Hints
- The input distribution is called P.
- The candidates and returned distribution are called Q.
- The rule evaluates the function associated with the bound.
- The selected posterior minimizes that function.
A description says: “Regularized risk minimization chooses a posterior using only its empirical loss.” Identify the missing component and explain its role.
Hints
- The missing component compares two distributions.
- One of those distributions is the prior P and the other is the posterior Q.
- This component is combined with empirical loss in the objective.
Essential Takeaways
- The PAC-Bayes bound can be converted from a theoretical guarantee into a learning rule.
- The rule begins with a prior P, considers possible posterior distributions Q, and returns a Q that minimizes the specified function.
- Regularized risk minimization combines empirical loss with the Kullback-Leibler distance between Q and P.
- The two quantities jointly minimized are empirical loss and KL distance.
- The selected posterior reflects the combined objective rather than either quantity considered alone.
Key Takeaways
- The PAC-Bayes bound leads to an optimization procedure, not only a statement about performance.
- The prior P is supplied to the rule, while the posterior Q is selected and returned.
- Regularized risk minimization combines empirical loss with the KL distance between Q and P.
- The learning rule jointly minimizes the data-fitting component and the distribution-comparison component.