Concepts / Marginal Distribution

Marginal Distribution

Realizability demands perfect agreement between some hypothesis in H and a target labeling function.

  • Programming

When Perfect Labels Fail

A traditional learning model often assumes that the features of every input determine one correct label with complete certainty. Under this assumption, learning means finding a hypothesis that agrees perfectly with a true labeling function. This requirement is called the realizability assumption.

matchesdefinesHypothesis hin HPerfect agreementevery domain pointTarget labelingfunctiontrue label for each point
What does it mean for one hypothesis in H to agree perfectly with the target labeling function on every domain point?

From Inputs and Labels to Pairs

Agnostic PAC learning relaxes the realizability requirement. Instead of requiring one perfectly correct labeling function, it describes how domain points and labels are generated together. The model uses a joint data-labels distribution over X × Y. Conceptually, an outcome from this distribution is a pair: a domain point together with a label.

combined withcombined withassigned probability byDomain pointfeature descriptionInput-label pairone outcome in X × YJoint distributiongenerates pairsLabeloutcome
How does a joint distribution assign probability to pairs consisting of a domain point and a label?

A generated papaya observation

Separate the two parts of an outcome involving papaya features and a taste label.

Domain point: The papaya's feature description is the domain point. It represents the input information available to the learning problem.

Label: The taste assessment is the label associated with that feature description.

Joint outcome: The feature description and taste label together form one input-label pair. A data-labels generating distribution describes probabilities for such pairs.

The joint view keeps the domain point and its label together instead of describing only inputs or only labels.

Frequency and Label Probability

The joint distribution contains two related but different views. The marginal distribution over domain points asks which feature descriptions are likely to appear. The conditional probability over labels asks how likely each label is once a particular feature description has been given.

describes frequency ofdescribes likelihood ofMarginaldistributionwhich points appearDomain pointfeature descriptionConditionalprobabilitylabels given a pointPossible labelsoutcomes for that point
What is the difference between how frequently a domain point occurs and the probabilities of its possible labels given that point?
QuestionDistributional viewWhat it describes
Which feature descriptions are likely to appear?Marginal distributionHow frequently domain points occur
Which labels are likely after a feature description is given?Conditional probabilityLabel uncertainty for that domain point
How likely is a particular feature-label pair?Joint distributionThe combined occurrence of a domain point and a label

The three views describe related parts of the same data-labels generating model.

One Point, Multiple Labels

A data-labels generating distribution can assign different labels to the same domain point. This is the key way the agnostic view represents situations in which the measured features do not completely determine the label. The same feature description can be paired with different outcomes, each with its own conditional probability.

can occur withcan occur withcan occur withFeature descriptionPone domain pointLabel Apossible outcomeLabel Bpossible outcomeLabel Cpossible outcome
How can the same domain point be associated with different labels under a data-labels generating distribution?

Separating occurrence from outcome

Suppose a particular papaya feature description appears in the data. Explain what the marginal and conditional views say about it.

Marginal question: Ask how likely that feature description is to appear among domain points. This is information from the marginal distribution.

Conditional question: After fixing that feature description, ask how likely each taste label is. This is information from the conditional probability over labels.

Multiple outcomes: If more than one taste label can be paired with the same feature description, the data-labels distribution represents that uncertainty directly.

The feature description can have one occurrence frequency while its possible labels have their own conditional probabilities.

Realizable and Nonrealizable Data

Under realizability, some hypothesis in H agrees perfectly with the target labeling function on every domain point. That means each domain point has a target label that the successful hypothesis can match without exception.

When measured features do not completely determine labels, the perfect-agreement requirement may be unrealistic. A single domain point may be associated with different labels in the data-labels distribution. In that setting, the learning problem is described through the joint distribution and its marginal and conditional parts rather than through one perfectly correct labeling function.

maps tocan be paired withDomain pointone target labelTarget labelperfect agreement possibleDomain pointsame feature descriptionMultiple labelsconditional uncertainty
What changes when labels are noisy or conflicting, so that no single hypothesis can perfectly match all observed labels?
  • Assuming every practical learning problem has one perfectly correct labeling function.

    Many practical problems do not satisfy realizability because measured features may not completely determine labels.

    Fix: Use the joint data-labels distribution when the relationship between features and labels can include uncertainty.

  • Calling the conditional label probability the marginal distribution.

    The marginal distribution concerns which domain points are likely to appear, while the conditional probability concerns labels given a domain point.

    Fix: First identify whether the question is about input occurrence or label likelihood after the input is fixed.

  • Assuming a domain point can be paired with only one label in the agnostic model.

    The data-labels distribution is specifically able to represent different labels for the same domain point.

    Fix: Represent the alternative labels through their conditional probabilities.

Recovering the Input Distribution

The joint distribution describes complete domain-point and label pairs. To focus only on domain points, combine the probabilities of all label outcomes associated with each domain point. This produces the marginal distribution over domain points. The labels are no longer being distinguished in that view; their outcomes have been combined.

includeincludeincludeyieldsPoint P with LabelAjoint outcomeCombine labeloutcomesfocus on Point PPoint Pmarginal occurrencePoint P with LabelBjoint outcomePoint P with LabelCjoint outcome
How are label outcomes combined to obtain the distribution over domain points alone?

Tracing one domain point

A joint distribution contains several pairs that use the same domain point P but different labels. Determine what remains when the question changes from complete pairs to domain points alone.

Start with joint outcomes: Keep each pair distinct while asking about the complete data-labels distribution.

Group by domain point: Place together all pairs whose domain component is P, regardless of which label they contain.

Combine label outcomes: Treat the grouped label possibilities as one combined contribution for the occurrence of P.

Read the marginal view: The result describes how likely P is as a domain point, without retaining separate label outcomes.

The marginal distribution over P is obtained from the joint distribution by combining the label-specific outcomes associated with P.

When analyzing a learning problem, name the question before choosing the distributional view. Use the joint distribution for complete input-label pairs, the marginal distribution for how domain points occur, and the conditional probability for label likelihood after a domain point is given.

Practice the Distinction

EASY

For a fixed feature description, decide which distributional view answers each question: How likely is this feature description to appear? How likely is each possible label after the feature description is given? How likely is the complete feature-description and label pair?

Hints
  • The first question ignores the label and focuses on domain-point occurrence.
  • The second fixes a domain point and asks about labels.
  • The third keeps both parts of the outcome together.
  1. To solve the practice task, match each question to its object: marginal distribution for domain-point frequency, conditional probability for labels given a domain point, and joint distribution for complete domain-point and label pairs.

Key Takeaways

  • Realizability requires some hypothesis in H to agree perfectly with a target labeling function on every domain point.
  • Realizability can be unrealistic when measured features do not completely determine labels.
  • Agnostic PAC learning uses a joint data-labels distribution over domain points and labels.
  • The marginal distribution describes which domain points are likely to appear.
  • The conditional probability describes how likely each label is after a particular domain point is given, and different labels may be associated with the same domain point.

Key Takeaways

  • Realizability assumes perfect agreement between a hypothesis in H and a target labeling function.
  • Practical features may not completely determine labels, so a perfect labeling function may not exist.
  • A joint data-labels distribution describes complete pairs of domain points and labels.
  • The marginal distribution answers how frequently domain points occur, while conditional probabilities describe labels given a domain point.
  • A single domain point can be paired with different labels under the data-labels generating distribution.