Concepts / Realizability Assumption

Realizability Assumption

Realizability demands perfect agreement between some hypothesis in H and a target labeling function.

  • Programming

When Perfect Prediction Fails

A traditional learning model may assume that the features of every input determine one correct label with complete certainty. Under that assumption, learning means finding a hypothesis that agrees perfectly with a true labeling function. This requirement is called the realizability assumption. It is useful as an idealized model, but many practical problems do not satisfy it because measured features may not completely determine the labels.

Tracing a Realizable Labeling

Imagine a domain X containing feature descriptions and a label set Y. A target labeling function assigns one correct label to each domain point. The realizability assumption says that the hypothesis class H contains at least one hypothesis that makes exactly the same assignment as this target labeling function everywhere in X.

perfect agreementTarget labelingfunctionx1 → y1; x2 → y2Hypothesis hx1 → y1; x2 → y2
What must be true for one hypothesis in H to agree perfectly with the target labeling function on every domain point?

Checking realizability

Suppose the target labeling function assigns label y1 to x1 and label y2 to x2. Hypothesis h makes those same two assignments.

Compare x1: The target assigns y1 to x1, and h also assigns y1 to x1.

Compare x2: The target assigns y2 to x2, and h also assigns y2 to x2.

Check the whole domain: Realizability requires this agreement for every domain point, not only for the points examined in a small sample.

If h agrees with the target labeling function on every domain point, the realizability requirement is satisfied.

Joint Data and Label Generation

Agnostic PAC learning relaxes the requirement for one perfectly correct labeling function. It uses a joint data-labels distribution over X × Y. Each outcome in this joint distribution is a complete pair: a domain point together with a label. The distribution describes how domain points and labels are generated together.

included pairincluded pairincluded pair(x1, y1)joint probabilityData-labelsdistributionover X × Y(x1, y2)joint probability(x2, y1)joint probability
How are probabilities assigned to complete domain-point and label pairs?

The joint viewpoint is more flexible than a single deterministic labeling rule. It can describe how likely a complete pair is, including pairs that use the same domain point with different labels. Thus, the data-labels distribution can represent uncertainty or ambiguity in the relationship between measured features and labels.

paired withpaired withxone domain pointy1probability for (x, y1)y2probability for (x, y2)
How can one domain point be paired with different labels, each with its own probability?

Marginal Versus Conditional Information

The joint distribution contains two related but different views. The marginal distribution over domain points asks which feature descriptions are likely to appear. The conditional probability over labels asks how likely each label is once a particular feature description has been given. The marginal focuses on the occurrence of the input; the conditional focuses on the labels associated with that input.

together describe the joint distributionMarginal over XHow often does x occur?Conditional over YWhich labels follow x?
What is the difference between how often a domain point occurs and the probability of each label given that point?

Consider domain points describing papaya features and labels describing taste. The marginal distribution concerns which papaya feature descriptions are likely to appear. After one feature description is fixed, the conditional probability concerns how likely each taste label is for that description. These are different questions, even though both are parts of the same joint data-labels distribution.

ViewQuestion answeredWhat it describes
Marginal distribution over domain pointsWhich feature descriptions are likely to appear?The occurrence of domain points
Conditional probability over labelsGiven a particular feature description, how likely is each label?The label distribution for that domain point
Joint data-labels distributionHow likely is a complete domain-point and label pair?The combined generation of domain points and labels

The marginal and conditional views describe different aspects of the same data-labels distribution.

From Deterministic Labels to Ambiguity

Under realizability, each domain point has one correct target label, and some hypothesis in H matches that target labeling function everywhere. In a practical problem, measured features may omit information that affects the label. The same measured domain point can then be associated with different labels in the data-labels distribution. When that happens, no single deterministic labeling function based only on those measured features is guaranteed to agree with every observed label.

determinescan be paired withcan be paired withxone target labelxsame measured featuresy1deterministic labely1one possible labely2another possible label
What changes when labels contain noise or ambiguity, so that no single hypothesis can perfectly match every observed label?

Common Interpretation Errors

  • Treating realizability as the claim that every practical dataset has one perfectly predictable label.

    Many practical problems do not satisfy realizability because measured features may not completely determine labels.

    Fix: Treat realizability as an assumption of the traditional learning model, and recognize that agnostic PAC learning relaxes it.

  • Confusing the marginal distribution with the conditional label probability.

    The marginal asks which domain points are likely to appear, while the conditional probability asks how likely each label is after a domain point is given.

    Fix: Identify whether the question concerns the occurrence of x or the label distribution given x.

  • Assuming that one domain point cannot be paired with more than one label.

    A data-labels distribution over X × Y can assign different labels to the same domain point, each with its own probability.

    Fix: Think in terms of complete domain-point and label pairs and their joint probabilities.

Check Your Understanding

MEDIUM

A learning problem uses measured feature descriptions as domain points. The same feature description can appear with two different labels. Which statement best describes the situation?

Hints
  • Separate the question of how often the feature description appears from the question of which labels it receives.
  • Ask whether one hypothesis can agree perfectly with every observed label for that domain point.
  • Recall that the joint data-labels distribution is defined over complete pairs in X × Y.

What do you think happens?

If a joint data-labels distribution assigns positive probability to both (x, y1) and (x, y2), can a deterministic target labeling function that assigns only one label to x agree with every generated pair?

  • Yes, because the marginal distribution chooses x
  • Yes, because both labels are part of the same hypothesis
  • No, because the same x is paired with different labels
  • No, because joint distributions cannot contain repeated domain points
Reveal answer

Answer: No, because the same x is paired with different labels.

A deterministic target labeling function assigns one label to x, so it cannot agree with both different labels paired with x. This is the kind of situation that motivates the more flexible data-labels distribution used in agnostic PAC learning.

Key Takeaways

  1. Realizability requires some hypothesis in H to agree perfectly with a target labeling function on every domain point.
  2. The assumption can be unrealistic when measured features do not completely determine labels.
  3. Agnostic PAC learning uses a joint data-labels distribution over X × Y instead of requiring one perfectly correct labeling function.
  4. The marginal distribution describes which domain points are likely to appear.
  5. The conditional probability describes how likely each label is given a particular domain point, and different labels may be paired with the same domain point.

Key Takeaways

  • Realizability is the assumption that a hypothesis in H can match the target labeling function everywhere.
  • Practical measured features may leave out information needed to determine one label, making perfect agreement unrealistic.
  • A data-labels generating distribution is a joint distribution over complete pairs of domain points and labels.
  • The marginal distribution concerns how often domain points occur, while the conditional probability concerns labels given a domain point.
  • The joint distribution can pair one domain point with different labels, each with its own probability.