Supervised Learning and Hypothesis Classes
PAC-Bayes bounds define a hierarchy over a hypothesis class H.
One Class, Many Possible Predictors
In supervised learning, it is tempting to imagine that a learning algorithm must return one final hypothesis. The PAC-Bayes view uses a broader perspective: a hypothesis class H contains possible hypotheses, and the learning process can describe its result as a distribution over that class. The central movement is from a prior distribution P over hypotheses before learning to a posterior distribution Q after learning.
The Class H
A hypothesis class H is the collection of hypotheses under consideration. An individual hypothesis h is one member of that class. The PAC-Bayes description does not focus only on choosing one member. Instead, it uses distributions over H to describe how hypotheses are organized before learning and how the learning algorithm represents its result afterward.
Do not confuse the class H with one hypothesis. H is the whole collection; h denotes an individual member selected or weighted by a distribution.
Before and After Learning
| Distribution | Timing | Role |
|---|---|---|
| P | Before learning | Assigns a probability or density P(h) to each hypothesis |
| Q | After the learning process | Describes the distribution produced by the learning algorithm over H |
The prior distribution P assigns a probability or density P(h) to each hypothesis before learning. It provides the initial distributional view of the hypothesis class. The posterior probability Q is the distribution over H produced by the learning algorithm. The important distinction is not simply that the symbols are different. P describes the hierarchy before learning, while Q describes the distribution resulting from the learning process.
From P to Q
A Distributional Learning Trace
Suppose H contains three hypotheses: h1, h2, and h3. Trace the roles of P and Q without assuming that the algorithm returns only one hypothesis.
Start with H: The hypothesis class contains h1, h2, and h3. These are possible hypotheses under consideration.
Assign the prior: Before learning, P assigns a probability or density to each hypothesis in H. This is the initial hierarchy over the class.
Run the learning process: The learning algorithm uses the distributional view of H and produces a posterior probability Q over the class.
Use Q: The result is not required to be one returned hypothesis. Q can be used to select hypotheses according to its distribution.
Learning changes the distribution used to select hypotheses: the prior P describes the before-learning view, and the posterior Q describes the distribution produced after learning.
The important change is therefore distributional. The hypothesis class H remains the space of hypotheses being considered, while the distribution used to describe or select them changes from P to Q through learning. This is why the PAC-Bayes view can describe learning without requiring a single final hypothesis.
Randomized Prediction
A posterior probability Q defines a randomized prediction rule. For an input x, the rule first selects a hypothesis h according to Q. It then predicts h(x). The prediction rule is randomized because the selected hypothesis comes from a distribution rather than being fixed in advance as one single hypothesis.
One Prediction Under Q
Let Q be a posterior distribution over a hypothesis class H containing h1, h2, and h3. What does the randomized prediction rule do for an input x?
Select: The rule selects one hypothesis h according to Q. The selected hypothesis could be any member supported by the posterior distribution.
Evaluate: The selected hypothesis is applied to the input x.
Predict: The rule outputs h(x), the prediction made by the selected hypothesis.
Q specifies how the hypothesis is selected; the selected hypothesis specifies the prediction h(x).
Mistakes About P, Q, and H
Treating H as if it were one hypothesis
H is the hypothesis class, the collection of hypotheses under consideration.
Fix:
Distinguish the class H from an individual hypothesis h selected from that class.Calling P the result of learning
P is the prior distribution assigned before learning.
Fix:
Use Q for the posterior distribution produced by the learning algorithm.Assuming the algorithm must return one hypothesis
In the PAC-Bayes view, the algorithm can work with a distribution over H rather than one hypothesis.
Fix:
Describe the result as a posterior probability Q over the hypothesis class when appropriate.Treating Q as a prediction by itself
Q defines the rule for selecting h; the selected hypothesis then produces h(x).
Fix:
Trace both stages: select h according to Q, then predict h(x).
Check the Distributional Trace
A learning algorithm considers a hypothesis class H and produces a posterior probability Q. Explain the complete prediction process for an input x, and identify which distribution describes the system before learning.
Hints
- First identify the role of P and its timing.
- Then state what Q selects.
- Finish by writing the prediction produced by the selected hypothesis.
What do you think happens?
If a PAC-Bayes learning algorithm is described by a posterior Q, must it return exactly one fixed hypothesis?
Reveal answer
Answer: No, it can represent its result as a distribution over H
The PAC-Bayes view allows the learning algorithm to produce a posterior probability Q over the hypothesis class rather than necessarily returning one hypothesis.
Summary
- H is a hypothesis class containing individual hypotheses h.
- The prior P assigns a probability or density to hypotheses before learning.
- The posterior Q is the distribution over H produced by the learning process.
- Q defines a randomized prediction rule: select h according to Q and predict h(x).
- PAC-Bayes bounds connect the prior and posterior distributional views instead of focusing only on one selected hypothesis.
Key Takeaways
- A hypothesis class H contains the possible hypotheses under consideration.
- P describes the prior hierarchy over hypotheses before learning.
- Q describes the distribution produced by learning over the same hypothesis class.
- A randomized prediction rule selects h according to Q and then predicts h(x).
- The PAC-Bayes perspective tracks how learning changes the hypothesis-selection distribution from P to Q.