Concepts / PAC Learning and Sample Complexity

PAC Learning and Sample Complexity

Uniform convergence connects a finite i.i.d. sample with representativeness of an underlying distribution.

  • Programming

From Finite Samples to Distribution Guarantees

A learning algorithm receives a finite sample, but the learning model is described relative to an underlying probability distribution. This creates a central question: when is the finite sample representative enough that conclusions drawn from it apply to the distribution? Uniform convergence formalizes this question for an entire hypothesis class, rather than for only one selected hypothesis.

Uniform convergence is a high-probability statement about all hypotheses in a class H at once. It asks whether their empirical errors are close to their true errors under D.

measured on samplemeasured under Dcomparecompareh in Heach hypothesisEmpirical errorfinite sampleTrue errordistribution DDifference at mostepsilonfor every h in H
How do the empirical errors of all hypotheses in H compare with their true errors under D?

The Uniform Convergence Guarantee

A hypothesis class H has the uniform convergence property for accuracy epsilon, confidence delta, and distribution D when a sample of size m is representative enough that, with probability at least 1 minus delta, the empirical error of every hypothesis in H is within epsilon of its true error under D.

The word uniform is important. The guarantee is not limited to the single hypothesis eventually chosen by a learning algorithm. It covers every hypothesis in H simultaneously. The probability statement also matters: the guarantee holds with probability at least 1 minus delta over the random drawing of the finite sample. It does not claim that every possible sample has the desired representativeness.

Reading a Hypothetical Guarantee

Suppose a statement says that H has uniform convergence for accuracy epsilon = 0.1, confidence parameter delta = 0.05, distribution D, and sample size m.

Accuracy: The empirical and true errors for every hypothesis in H must be within 0.1 of one another.

Confidence: The statement is required to hold with probability at least 0.95, because 1 minus delta equals 0.95.

Distribution: True error is interpreted relative to the underlying distribution D.

Scope: The comparison applies across the whole class H, not just one hypothesis.

The guarantee says that a randomly drawn finite sample gives uniformly representative error estimates with at least 0.95 probability.

Four Parameters with Different Jobs

A uniform-convergence statement connects four pieces: epsilon, delta, D, and m. Each describes a different part of the guarantee. Keeping their roles separate prevents the common mistake of treating sample size, accuracy, and confidence as interchangeable.

sets allowed error gapsets failure probabilitydefines true errorprovides finite sampleepsilonaccuracy tolerancedeltafailure allowanceDunderlying distributionmsample sizeUniform convergencefor H
How are accuracy, confidence, the distribution, and sample size connected in a uniform-convergence guarantee?
SymbolRole in the guarantee
epsilonThe permitted difference between empirical and true error.
deltaThe amount subtracted from 1 to express the allowed failure probability.
DThe underlying probability distribution used to define true error.
mThe number of independently drawn examples in the finite sample.

Interpret the parameters by asking what each one controls.

The Sample-Size Threshold

The function m_UC_H(epsilon, delta) denotes the minimal sample complexity for uniform convergence for the hypothesis class H at the selected accuracy and confidence parameters. It is a threshold: the defining sample-size condition is m at least m_UC_H(epsilon, delta). When that condition is met, the uniform-convergence guarantee applies with probability at least 1 minus delta.

condition checkedcondition satisfiedm below thresholdm less than m_UC_Hm at thresholdm at least m_UC_HNo stated guaranteefrom this condition aloneUniform convergenceprobability at least 1minus delta
What sample-size threshold does m_UC_H provide, and how does increasing m move the problem into the guaranteed region?

Using a Hypothetical Threshold

Assume that m_UC_H(epsilon, delta) equals 500 for the selected H, epsilon, delta, and D. Compare two proposed sample sizes: 420 and 500.

First sample: For m = 420, the condition m at least m_UC_H(epsilon, delta) is false because 420 is below 500.

Threshold sample: For m = 500, the condition is true because the sample size reaches the stated minimum.

Larger samples: Any sample size greater than 500 also satisfies the same threshold condition.

The hypothetical uniform-convergence guarantee is supported by m = 500, but not by m = 420.

Testing a Proposed Sample Size

compareat leastgreater than proposedProposed msample sizem_UC_Hminimum sample sizeCondition satisfiedm at least m_UC_HCondition notsatisfiedm below m_UC_H
How can a proposed sample size be compared with m_UC_H(epsilon, delta) to determine whether the guarantee holds?
  1. Identify the relevant hypothesis class H and the selected epsilon and delta.
  2. Read the corresponding threshold m_UC_H(epsilon, delta).
  3. Compare the proposed sample size m with that threshold.
  4. If m is at least the threshold, the stated uniform-convergence guarantee applies with probability at least 1 minus delta.
  5. If m is below the threshold, the given condition does not establish that guarantee.
EASY

Suppose m_UC_H(epsilon, delta) is 800. Decide whether each proposed sample size satisfies the defining condition: m = 799, m = 800, and m = 950. For each one, state whether the uniform-convergence guarantee is supported by the threshold condition.

Hints
  • Compare each proposed value directly with 800.
  • Equality satisfies the condition m at least m_UC_H(epsilon, delta).
  • A value below the threshold does not satisfy the stated condition.

Common Interpretation Errors

  • Treating uniform convergence as a claim about only the selected hypothesis.

    The defining comparison applies to every hypothesis in H simultaneously.

    Fix: Keep the class-wide scope visible: the guarantee concerns all hypotheses in H.

  • Reading probability at least 1 minus delta as a guarantee for every possible sample.

    The guarantee holds with high probability over the random sample, not necessarily for every possible sample.

    Fix: Interpret delta as allowing a failure probability of delta.

  • Confusing epsilon with delta.

    Epsilon controls the accuracy gap, while delta controls the confidence level.

    Fix: Associate epsilon with closeness and delta with the probability statement.

  • Using a sample size below the threshold as if it proved uniform convergence.

    The defining condition requires m to be at least m_UC_H(epsilon, delta).

    Fix: Compare the proposed m directly with the threshold before making the conclusion.

  • Ignoring the distribution D.

    The learning model is described relative to an underlying probability distribution, and true error is interpreted under D.

    Fix: Include D when identifying what the empirical sample is meant to represent.

Uniform Convergence Checklist

  1. Uniform convergence connects a finite i.i.d. sample with the representativeness of an underlying distribution.
  2. For every hypothesis in H, empirical error should be within epsilon of true error under D, with probability at least 1 minus delta.
  3. Epsilon controls accuracy, delta controls confidence, D supplies the underlying distribution, and m is the finite sample size.
  4. The function m_UC_H(epsilon, delta) gives the minimal sample complexity for the guarantee.
  5. The defining test is m at least m_UC_H(epsilon, delta).

Key Takeaways

  • Uniform convergence asks whether one finite sample gives representative error estimates for every hypothesis in H.
  • The guarantee compares empirical error with true error under D, allowing an accuracy gap of epsilon.
  • The probability statement is at least 1 minus delta, so it is a high-probability claim rather than a claim about every possible sample.
  • m_UC_H(epsilon, delta) is the minimum sample-size threshold, and the condition m at least that threshold supports the guarantee.