Conditional Probability
Realizability demands perfect agreement between some hypothesis in H and a target labeling function.
When One Label Is Not Enough
A traditional learning model may assume that the features of every input determine one correct label with complete certainty. Under that assumption, learning means finding a hypothesis that agrees perfectly with a true labeling function. This is called the realizability assumption. The difficulty is that measured features may not completely determine labels in many practical problems. The same feature description can therefore be associated with different labels in the data-labels model.
From Inputs to Labeled Examples
Agnostic PAC learning relaxes realizability. Instead of requiring one perfectly correct labeling function, it uses a joint data-labels distribution over X × Y. Here, X represents domain points and Y represents labels. A draw from this joint distribution produces a combination: one domain point together with one label. The distribution describes how these combinations are generated together.
Papaya Features and Taste Labels
Separate the two questions represented by a data-labels distribution over papaya examples.
Identify the domain point: A papaya's feature description is a domain point in X.
Identify the label: The papaya's taste label is an element of Y.
View the pair: The feature description and taste label together form one possible combination in the joint distribution.
Separate the questions: The marginal distribution asks which papaya feature descriptions are likely to appear. The conditional probability asks how likely each taste label is after a particular feature description is given.
The marginal distribution describes the inputs, while the conditional probability describes label likelihoods for a selected input. Together they describe the joint distribution over papaya features and taste labels.
Marginal Versus Conditional
The joint data-labels distribution contains two connected views. The marginal distribution over domain points describes which feature descriptions are likely to appear, without focusing on one particular label. The conditional probability over labels answers a different question: once a particular domain point or feature description is given, how likely is each label? The first view concerns the distribution of inputs; the second concerns the labels associated with a selected input.
Repeated Inputs with Different Labels
The joint distribution does not require every domain point to have only one possible label. A particular feature description can be paired with different labels. This is how the data-labels model represents situations in which measured features do not completely determine the label. The conditional probability over labels records the likelihood of the possible labels after that feature description is given.
Reading an Ambiguous Feature Description
Suppose one papaya feature description appears in the domain. What does the data-labels distribution need to represent if that description can occur with more than one taste label?
Hold the input fixed: Focus on one particular papaya feature description rather than changing the input.
Allow multiple pairings: The same feature description can be paired with one taste label in some generated examples and another taste label in other generated examples.
Interpret the conditional view: The conditional probability over labels describes how likely each taste label is for that fixed feature description.
Interpret the joint view: The joint distribution contains the corresponding combinations of that feature description and each possible label.
A fixed domain point does not force a single label in the data-labels model. Different labels can be represented through different input-label combinations and their associated probabilities.
Common Reasoning Errors
Assuming that every domain point must have one certain label.
The agnostic data-labels model permits different labels for the same domain point when measured features do not completely determine labels.
Fix:
Use the conditional probability over labels to describe how likely each label is for the selected domain point.Confusing the marginal distribution with the conditional probability.
The marginal distribution concerns which domain points or feature descriptions are likely to appear, while the conditional probability concerns labels after a particular domain point is given.
Fix:
First identify whether the question is about the inputs that occur or about labels associated with one selected input.Treating realizability as a requirement of every learning problem.
Many practical problems do not satisfy realizability because measured features may not completely determine labels.
Fix:
Recognize that agnostic PAC learning relaxes this requirement by modeling a joint distribution over domain points and labels.
Check Your Understanding
A learning problem records feature descriptions and labels. The same feature description can occur with two different labels. Explain what the marginal distribution describes, what the conditional probability describes, and why this situation does not satisfy the strict realizability assumption.
Hints
- Start with the question: which domain points are likely to appear?
- Then ask: after one domain point is given, what does the conditional probability describe?
- Realizability requires perfect agreement between some hypothesis in H and a target labeling function.
Practice Answer
Explain the three parts of the situation.
Marginal distribution: It describes which feature descriptions, or domain points, are likely to appear.
Conditional probability: It describes how likely each label is once a particular feature description is given.
Realizability: If one feature description can occur with different labels, the features do not completely determine one correct label in the modeled data. The strict perfect-agreement assumption is therefore not appropriate for that situation.
The joint distribution keeps both views together: it describes the occurrence of domain points and the label probabilities associated with each one.
Key Takeaways
- Realizability requires perfect agreement between some hypothesis in H and a target labeling function.
- Measured features may not completely determine labels, so many practical problems do not satisfy realizability.
- Agnostic PAC learning uses a joint data-labels distribution over domain points and labels.
- The marginal distribution describes which domain points are likely to appear.
- The conditional probability describes how likely each label is after a particular domain point is given, and it permits different labels for the same domain point.
Key Takeaways
- Realizability assumes that one hypothesis in H agrees perfectly with a target labeling function.
- Practical learning problems may violate this assumption because measured features do not always completely determine labels.
- A joint data-labels distribution describes domain points and labels together.
- The marginal distribution concerns which inputs appear, while conditional probability concerns label likelihoods for a selected input.
- The same domain point can be paired with different labels in the agnostic model.