Concepts / Independent Sampling

Independent Sampling

D is the distribution; z ∼ D describes sampling from it.

  • Programming

From Distribution to Sample

Probability notation becomes easier when you read it from the distribution outward. Start with D: it names a probability distribution over a set Z. Then read z ∼ D as a statement about the relationship between z and D. It says that z is sampled according to D.

The symbol z ∼ D does not identify one permanently fixed value of z. It tells you to regard z as a value obtained by sampling from D.

sample according toDdistributionzsampled value
What happens when a value z is sampled from the distribution D, and how is that different from z being fixed?

Using a Sampled Variable

Once z is understood as sampled from D, later expressions involving z are evaluated with respect to that sampling convention. For example, a function can receive z as its input, or a condition can be evaluated for the sampled value. The notation keeps track of the distribution that governs the variable.

The expression E z ∼ D [f(z)] denotes the expected value of f(z) when z is sampled according to D. The subscript identifies the distribution governing the variable inside the brackets.

Following the sampling context

Interpret the expression E z ∼ D [f(z)].

Identify D: D is the distribution over the underlying set Z.

Read the subscript: The notation z ∼ D says that z is sampled according to D.

Read the brackets: The function f is evaluated using the sampled variable z.

Interpret the whole expression: The expression summarizes the expected value of the real-valued quantity f(z) under that sampling process.

It is an expected value governed by D, not a statement that z is one fixed value.

Averages and Event Probabilities

Expected-value notation and probability notation both refer to sampling from D, but they summarize different things. Expected value summarizes a real-valued function of the sampled variable. Probability notation summarizes whether a Boolean condition is true.

NotationWhat is evaluatedWhat is summarized
E z ∼ D [f(z)]A real-valued function f(z)The expected value of that quantity
P z ∼ D [f(z)]A Boolean condition f(z)The probability that the condition is true
E z ∼ D [f(z)]real-valued quantityP z ∼ D [f(z)]Boolean condition
What is the difference between averaging a numerical quantity over samples and measuring the probability that an event occurs?

If f maps Z to {true, false}, then P z ∼ D [f(z)] denotes the probability that f(z) is true when z is sampled according to D. In set notation, this is the probability assigned by D to the collection of values z for which f(z) is true.

Building the Product Distribution

A single sampled variable is not the only object that can have a distribution. Suppose m points are sampled from D. The resulting object is an m-component tuple written as (z₁, …, zₘ). Each zᵢ is sampled from D independently of the other points.

D^m refers to the probability over Z^m induced by sampling m points from D independently. D governs individual samples from Z; D^m governs the larger space Z^m of m-component sample tuples.

sample independentlysample independentlysample independentlycomponentcomponentcomponentinducesDdistribution over Zz₁sample from D(z₁, …, zₘ)m-component tupleD^mprobability over Z^mz₂sample from Dzₘsample from D
How do m independently sampled points from D combine to form the product distribution D^m?

The superscript in D^m signals repeated independent sampling and the larger sample space Z^m. It does not describe a new individual sample variable.

Common Reading Errors

  • Treating z ∼ D as if z were a permanently fixed value.

    The notation describes z as a value obtained by sampling from D.

    Fix: Carry the sampling context into later expressions involving z.

  • Confusing an expected value with an event probability.

    Expected value summarizes a real-valued function, while probability notation summarizes when a Boolean condition is true.

    Fix: Ask whether the expression is averaging a numerical quantity or measuring the probability of a true condition.

  • Reading D^m as an individual sample from D.

    D^m refers to the probability over m-component tuples in Z^m.

    Fix: Look for the tuple (z₁, …, zₘ) and remember that its components are sampled independently from D.

  • Assuming the shorthand E[f] introduces a different expected value.

    The shorter notation omits sampling information that is already clear from context.

    Fix: Restore the distribution and sampled variable mentally when interpreting the shorthand.

Practice the Notation

MEDIUM

For each expression, identify what is being described: the distribution of one sample, an expected value, an event probability, or a distribution over tuples. Then explain the role of D.

Hints
  • Read D first as a distribution over Z.
  • Interpret z ∼ D as the sampling relationship.
  • For an expectation, look for a real-valued function.
  • For a probability, look for a Boolean condition.
  • For D^m, identify the m-component tuple and the space Z^m.

Classifying Four Notations

Classify the roles of D, z ∼ D, E z ∼ D [f(z)], P z ∼ D [f(z)], and D^m.

D: This names a probability distribution over the set Z.

z ∼ D: This states that z is sampled according to D.

E z ∼ D [f(z)]: This is the expected value of a real-valued function of the sampled variable.

P z ∼ D [f(z)]: This is the probability that the Boolean condition f(z) is true under the sampling convention.

D^m: This is the induced probability over m-component tuples in Z^m when the components are sampled independently from D.

The notation changes meaning according to the object being summarized: one distribution, one sampled variable, a numerical average, an event probability, or a tuple distribution.

Key Takeaways

  • D names a probability distribution over an underlying set Z.
  • z ∼ D says that z is sampled according to D rather than treated as one permanently fixed value.
  • E z ∼ D [f(z)] summarizes the expected value of a real-valued function, while P z ∼ D [f(z)] summarizes the probability that a Boolean condition is true.
  • D^m is the induced probability over m-component tuples in Z^m formed by sampling each component independently from D.
  • The superscript in D^m describes repeated independent sampling and a larger sample space, not a new individual sample variable.