Concepts / Supervised Learning and Hypothesis Classes

Supervised Learning and Hypothesis Classes

PAC-Bayes bounds define a hierarchy over a hypothesis class H.

  • Programming

One Class, Many Possible Predictors

In supervised learning, it is tempting to imagine that a learning algorithm must return one final hypothesis. The PAC-Bayes view uses a broader perspective: a hypothesis class H contains possible hypotheses, and the learning process can describe its result as a distribution over that class. The central movement is from a prior distribution P over hypotheses before learning to a posterior distribution Q after learning.

organizescontainscontainscontainsPrior Pdistribution beforelearningHhypothesis classh1a hypothesish2a hypothesish3a hypothesis
How does the PAC-Bayes perspective organize hypotheses inside H?

The Class H

A hypothesis class H is the collection of hypotheses under consideration. An individual hypothesis h is one member of that class. The PAC-Bayes description does not focus only on choosing one member. Instead, it uses distributions over H to describe how hypotheses are organized before learning and how the learning algorithm represents its result afterward.

containscontainscontainsHhypothesis classhAhypothesishBhypothesishChypothesis
What does H contain, and how do individual hypotheses relate to the class?

Do not confuse the class H with one hypothesis. H is the whole collection; h denotes an individual member selected or weighted by a distribution.

Before and After Learning

DistributionTimingRole
PBefore learningAssigns a probability or density P(h) to each hypothesis
QAfter the learning processDescribes the distribution produced by the learning algorithm over H

The prior distribution P assigns a probability or density P(h) to each hypothesis before learning. It provides the initial distributional view of the hypothesis class. The posterior probability Q is the distribution over H produced by the learning algorithm. The important distinction is not simply that the symbols are different. P describes the hierarchy before learning, while Q describes the distribution resulting from the learning process.

assignsassignsassignsassignsPbefore learningh1P(h1)h1Q(h1)h2P(h2)Qafter learningh2Q(h2)
What is different about the distribution used before learning and the distribution produced after learning?

From P to Q

provides initial distributionproducesPprior over HLearninglearning processQposterior over H
What changes as the learning process transforms the prior distribution into the posterior distribution?

A Distributional Learning Trace

Suppose H contains three hypotheses: h1, h2, and h3. Trace the roles of P and Q without assuming that the algorithm returns only one hypothesis.

Start with H: The hypothesis class contains h1, h2, and h3. These are possible hypotheses under consideration.

Assign the prior: Before learning, P assigns a probability or density to each hypothesis in H. This is the initial hierarchy over the class.

Run the learning process: The learning algorithm uses the distributional view of H and produces a posterior probability Q over the class.

Use Q: The result is not required to be one returned hypothesis. Q can be used to select hypotheses according to its distribution.

Learning changes the distribution used to select hypotheses: the prior P describes the before-learning view, and the posterior Q describes the distribution produced after learning.

The important change is therefore distributional. The hypothesis class H remains the space of hypotheses being considered, while the distribution used to describe or select them changes from P to Q through learning. This is why the PAC-Bayes view can describe learning without requiring a single final hypothesis.

Randomized Prediction

A posterior probability Q defines a randomized prediction rule. For an input x, the rule first selects a hypothesis h according to Q. It then predicts h(x). The prediction rule is randomized because the selected hypothesis comes from a distribution rather than being fixed in advance as one single hypothesis.

select according to Qreceivesapply hQposterior over Hhselected hypothesisxinputh(x)prediction
How does a posterior distribution select a hypothesis and use it to produce a prediction?

One Prediction Under Q

Let Q be a posterior distribution over a hypothesis class H containing h1, h2, and h3. What does the randomized prediction rule do for an input x?

Select: The rule selects one hypothesis h according to Q. The selected hypothesis could be any member supported by the posterior distribution.

Evaluate: The selected hypothesis is applied to the input x.

Predict: The rule outputs h(x), the prediction made by the selected hypothesis.

Q specifies how the hypothesis is selected; the selected hypothesis specifies the prediction h(x).

Mistakes About P, Q, and H

  • Treating H as if it were one hypothesis

    H is the hypothesis class, the collection of hypotheses under consideration.

    Fix: Distinguish the class H from an individual hypothesis h selected from that class.

  • Calling P the result of learning

    P is the prior distribution assigned before learning.

    Fix: Use Q for the posterior distribution produced by the learning algorithm.

  • Assuming the algorithm must return one hypothesis

    In the PAC-Bayes view, the algorithm can work with a distribution over H rather than one hypothesis.

    Fix: Describe the result as a posterior probability Q over the hypothesis class when appropriate.

  • Treating Q as a prediction by itself

    Q defines the rule for selecting h; the selected hypothesis then produces h(x).

    Fix: Trace both stages: select h according to Q, then predict h(x).

Check the Distributional Trace

MEDIUM

A learning algorithm considers a hypothesis class H and produces a posterior probability Q. Explain the complete prediction process for an input x, and identify which distribution describes the system before learning.

Hints
  • First identify the role of P and its timing.
  • Then state what Q selects.
  • Finish by writing the prediction produced by the selected hypothesis.

What do you think happens?

If a PAC-Bayes learning algorithm is described by a posterior Q, must it return exactly one fixed hypothesis?

  • Yes, exactly one fixed hypothesis
  • No, it can represent its result as a distribution over H
Reveal answer

Answer: No, it can represent its result as a distribution over H

The PAC-Bayes view allows the learning algorithm to produce a posterior probability Q over the hypothesis class rather than necessarily returning one hypothesis.

Summary

  1. H is a hypothesis class containing individual hypotheses h.
  2. The prior P assigns a probability or density to hypotheses before learning.
  3. The posterior Q is the distribution over H produced by the learning process.
  4. Q defines a randomized prediction rule: select h according to Q and predict h(x).
  5. PAC-Bayes bounds connect the prior and posterior distributional views instead of focusing only on one selected hypothesis.

Key Takeaways

  • A hypothesis class H contains the possible hypotheses under consideration.
  • P describes the prior hierarchy over hypotheses before learning.
  • Q describes the distribution produced by learning over the same hypothesis class.
  • A randomized prediction rule selects h according to Q and then predicts h(x).
  • The PAC-Bayes perspective tracks how learning changes the hypothesis-selection distribution from P to Q.