Marginal Distribution
Realizability demands perfect agreement between some hypothesis in H and a target labeling function.
When Perfect Labels Fail
A traditional learning model often assumes that the features of every input determine one correct label with complete certainty. Under this assumption, learning means finding a hypothesis that agrees perfectly with a true labeling function. This requirement is called the realizability assumption.
From Inputs and Labels to Pairs
Agnostic PAC learning relaxes the realizability requirement. Instead of requiring one perfectly correct labeling function, it describes how domain points and labels are generated together. The model uses a joint data-labels distribution over X × Y. Conceptually, an outcome from this distribution is a pair: a domain point together with a label.
A generated papaya observation
Separate the two parts of an outcome involving papaya features and a taste label.
Domain point: The papaya's feature description is the domain point. It represents the input information available to the learning problem.
Label: The taste assessment is the label associated with that feature description.
Joint outcome: The feature description and taste label together form one input-label pair. A data-labels generating distribution describes probabilities for such pairs.
The joint view keeps the domain point and its label together instead of describing only inputs or only labels.
Frequency and Label Probability
The joint distribution contains two related but different views. The marginal distribution over domain points asks which feature descriptions are likely to appear. The conditional probability over labels asks how likely each label is once a particular feature description has been given.
| Question | Distributional view | What it describes |
|---|---|---|
| Which feature descriptions are likely to appear? | Marginal distribution | How frequently domain points occur |
| Which labels are likely after a feature description is given? | Conditional probability | Label uncertainty for that domain point |
| How likely is a particular feature-label pair? | Joint distribution | The combined occurrence of a domain point and a label |
The three views describe related parts of the same data-labels generating model.
One Point, Multiple Labels
A data-labels generating distribution can assign different labels to the same domain point. This is the key way the agnostic view represents situations in which the measured features do not completely determine the label. The same feature description can be paired with different outcomes, each with its own conditional probability.
Separating occurrence from outcome
Suppose a particular papaya feature description appears in the data. Explain what the marginal and conditional views say about it.
Marginal question: Ask how likely that feature description is to appear among domain points. This is information from the marginal distribution.
Conditional question: After fixing that feature description, ask how likely each taste label is. This is information from the conditional probability over labels.
Multiple outcomes: If more than one taste label can be paired with the same feature description, the data-labels distribution represents that uncertainty directly.
The feature description can have one occurrence frequency while its possible labels have their own conditional probabilities.
Realizable and Nonrealizable Data
Under realizability, some hypothesis in H agrees perfectly with the target labeling function on every domain point. That means each domain point has a target label that the successful hypothesis can match without exception.
When measured features do not completely determine labels, the perfect-agreement requirement may be unrealistic. A single domain point may be associated with different labels in the data-labels distribution. In that setting, the learning problem is described through the joint distribution and its marginal and conditional parts rather than through one perfectly correct labeling function.
Assuming every practical learning problem has one perfectly correct labeling function.
Many practical problems do not satisfy realizability because measured features may not completely determine labels.
Fix:
Use the joint data-labels distribution when the relationship between features and labels can include uncertainty.Calling the conditional label probability the marginal distribution.
The marginal distribution concerns which domain points are likely to appear, while the conditional probability concerns labels given a domain point.
Fix:
First identify whether the question is about input occurrence or label likelihood after the input is fixed.Assuming a domain point can be paired with only one label in the agnostic model.
The data-labels distribution is specifically able to represent different labels for the same domain point.
Fix:
Represent the alternative labels through their conditional probabilities.
Recovering the Input Distribution
The joint distribution describes complete domain-point and label pairs. To focus only on domain points, combine the probabilities of all label outcomes associated with each domain point. This produces the marginal distribution over domain points. The labels are no longer being distinguished in that view; their outcomes have been combined.
Tracing one domain point
A joint distribution contains several pairs that use the same domain point P but different labels. Determine what remains when the question changes from complete pairs to domain points alone.
Start with joint outcomes: Keep each pair distinct while asking about the complete data-labels distribution.
Group by domain point: Place together all pairs whose domain component is P, regardless of which label they contain.
Combine label outcomes: Treat the grouped label possibilities as one combined contribution for the occurrence of P.
Read the marginal view: The result describes how likely P is as a domain point, without retaining separate label outcomes.
The marginal distribution over P is obtained from the joint distribution by combining the label-specific outcomes associated with P.
When analyzing a learning problem, name the question before choosing the distributional view. Use the joint distribution for complete input-label pairs, the marginal distribution for how domain points occur, and the conditional probability for label likelihood after a domain point is given.
Practice the Distinction
For a fixed feature description, decide which distributional view answers each question: How likely is this feature description to appear? How likely is each possible label after the feature description is given? How likely is the complete feature-description and label pair?
Hints
- The first question ignores the label and focuses on domain-point occurrence.
- The second fixes a domain point and asks about labels.
- The third keeps both parts of the outcome together.
- To solve the practice task, match each question to its object: marginal distribution for domain-point frequency, conditional probability for labels given a domain point, and joint distribution for complete domain-point and label pairs.
Key Takeaways
- Realizability requires some hypothesis in H to agree perfectly with a target labeling function on every domain point.
- Realizability can be unrealistic when measured features do not completely determine labels.
- Agnostic PAC learning uses a joint data-labels distribution over domain points and labels.
- The marginal distribution describes which domain points are likely to appear.
- The conditional probability describes how likely each label is after a particular domain point is given, and different labels may be associated with the same domain point.
Key Takeaways
- Realizability assumes perfect agreement between a hypothesis in H and a target labeling function.
- Practical features may not completely determine labels, so a perfect labeling function may not exist.
- A joint data-labels distribution describes complete pairs of domain points and labels.
- The marginal distribution answers how frequently domain points occur, while conditional probabilities describe labels given a domain point.
- A single domain point can be paired with different labels under the data-labels generating distribution.