Concepts / Conditional Probability

Conditional Probability

Realizability demands perfect agreement between some hypothesis in H and a target labeling function.

  • Programming

When One Label Is Not Enough

A traditional learning model may assume that the features of every input determine one correct label with complete certainty. Under that assumption, learning means finding a hypothesis that agrees perfectly with a true labeling function. This is called the realizability assumption. The difficulty is that measured features may not completely determine labels in many practical problems. The same feature description can therefore be associated with different labels in the data-labels model.

agrees withagrees withmay not determinecan varyHypothesis hmember of HPerfect agreementMeasured featuresmay not determine labelsNo perfect agreementTarget labelingfunctionone label per domain pointObserved labelscan vary
What does a perfectly realizable labeling look like compared with a practical setting where measured features may not completely determine labels?

From Inputs to Labeled Examples

Agnostic PAC learning relaxes realizability. Instead of requiring one perfectly correct labeling function, it uses a joint data-labels distribution over X × Y. Here, X represents domain points and Y represents labels. A draw from this joint distribution produces a combination: one domain point together with one label. The distribution describes how these combinations are generated together.

includesincludescombined withcombined withJoint distributionover X × YDomain pointxLabeled example(x, y)Labely
How does one joint distribution assign probability to combinations of domain points and labels?

Papaya Features and Taste Labels

Separate the two questions represented by a data-labels distribution over papaya examples.

Identify the domain point: A papaya's feature description is a domain point in X.

Identify the label: The papaya's taste label is an element of Y.

View the pair: The feature description and taste label together form one possible combination in the joint distribution.

Separate the questions: The marginal distribution asks which papaya feature descriptions are likely to appear. The conditional probability asks how likely each taste label is after a particular feature description is given.

The marginal distribution describes the inputs, while the conditional probability describes label likelihoods for a selected input. Together they describe the joint distribution over papaya features and taste labels.

Marginal Versus Conditional

The joint data-labels distribution contains two connected views. The marginal distribution over domain points describes which feature descriptions are likely to appear, without focusing on one particular label. The conditional probability over labels answers a different question: once a particular domain point or feature description is given, how likely is each label? The first view concerns the distribution of inputs; the second concerns the labels associated with a selected input.

view ofview ofdescribesdescribesJoint distributiondomain points and labelsMarginal distributionwhich inputs appearDomain pointsfeature descriptionsConditionalprobabilitywhich labels follow aninputLabelslabel likelihoods
What does each probability describe, and how are the distribution of inputs and the label probabilities connected?

Repeated Inputs with Different Labels

The joint distribution does not require every domain point to have only one possible label. A particular feature description can be paired with different labels. This is how the data-labels model represents situations in which measured features do not completely determine the label. The conditional probability over labels records the likelihood of the possible labels after that feature description is given.

can pair withcan pair withformsformsFeature descriptionsame domain pointTaste label Apossible labelInput-label pair A(same point, A)Taste label Bpossible labelInput-label pair B(same point, B)
How can the same domain point appear with different labels, and how is that represented in the distribution?

Reading an Ambiguous Feature Description

Suppose one papaya feature description appears in the domain. What does the data-labels distribution need to represent if that description can occur with more than one taste label?

Hold the input fixed: Focus on one particular papaya feature description rather than changing the input.

Allow multiple pairings: The same feature description can be paired with one taste label in some generated examples and another taste label in other generated examples.

Interpret the conditional view: The conditional probability over labels describes how likely each taste label is for that fixed feature description.

Interpret the joint view: The joint distribution contains the corresponding combinations of that feature description and each possible label.

A fixed domain point does not force a single label in the data-labels model. Different labels can be represented through different input-label combinations and their associated probabilities.

Common Reasoning Errors

  • Assuming that every domain point must have one certain label.

    The agnostic data-labels model permits different labels for the same domain point when measured features do not completely determine labels.

    Fix: Use the conditional probability over labels to describe how likely each label is for the selected domain point.

  • Confusing the marginal distribution with the conditional probability.

    The marginal distribution concerns which domain points or feature descriptions are likely to appear, while the conditional probability concerns labels after a particular domain point is given.

    Fix: First identify whether the question is about the inputs that occur or about labels associated with one selected input.

  • Treating realizability as a requirement of every learning problem.

    Many practical problems do not satisfy realizability because measured features may not completely determine labels.

    Fix: Recognize that agnostic PAC learning relaxes this requirement by modeling a joint distribution over domain points and labels.

Check Your Understanding

MEDIUM

A learning problem records feature descriptions and labels. The same feature description can occur with two different labels. Explain what the marginal distribution describes, what the conditional probability describes, and why this situation does not satisfy the strict realizability assumption.

Hints
  • Start with the question: which domain points are likely to appear?
  • Then ask: after one domain point is given, what does the conditional probability describe?
  • Realizability requires perfect agreement between some hypothesis in H and a target labeling function.

Practice Answer

Explain the three parts of the situation.

Marginal distribution: It describes which feature descriptions, or domain points, are likely to appear.

Conditional probability: It describes how likely each label is once a particular feature description is given.

Realizability: If one feature description can occur with different labels, the features do not completely determine one correct label in the modeled data. The strict perfect-agreement assumption is therefore not appropriate for that situation.

The joint distribution keeps both views together: it describes the occurrence of domain points and the label probabilities associated with each one.

Key Takeaways

  1. Realizability requires perfect agreement between some hypothesis in H and a target labeling function.
  2. Measured features may not completely determine labels, so many practical problems do not satisfy realizability.
  3. Agnostic PAC learning uses a joint data-labels distribution over domain points and labels.
  4. The marginal distribution describes which domain points are likely to appear.
  5. The conditional probability describes how likely each label is after a particular domain point is given, and it permits different labels for the same domain point.

Key Takeaways

  • Realizability assumes that one hypothesis in H agrees perfectly with a target labeling function.
  • Practical learning problems may violate this assumption because measured features do not always completely determine labels.
  • A joint data-labels distribution describes domain points and labels together.
  • The marginal distribution concerns which inputs appear, while conditional probability concerns label likelihoods for a selected input.
  • The same domain point can be paired with different labels in the agnostic model.