Concepts / Expected Values

Expected Values

D is the distribution; z ∼ D describes sampling from it.

  • Programming

From Distribution to Sample

Expected-value notation becomes much easier to read when you begin with the distribution. D names a probability distribution over a set. The expression z ∼ D then says that z is sampled according to D. Once that sampling relationship is clear, the notation inside an expectation or a probability statement tells you what is being evaluated for the sampled value.

assigns probabilityassigns probabilityassigns probabilityDdistributionz₁possible valuez₂possible valuez₃possible value
What does D contain, and how does it associate probabilities with the possible values in its set?

Reading z ∼ D

The notation z ∼ D describes a relationship between the variable z and the distribution D. It tells you to regard z as a value obtained by sampling from D. It does not identify one permanently fixed value of z. Instead, later expressions involving z are interpreted with respect to the sampling convention established by D.

samplesevaluatesDdistributionzsampled valuef(z)later evaluation
What happens when z ∼ D: how is a value z selected according to the probabilities in D?

What do you think happens?

When you read z ∼ D, is z one permanently fixed value or a value regarded as sampled from D?

  • One permanently fixed value
  • A value obtained by sampling from D
Reveal answer

Answer: A value obtained by sampling from D

The notation establishes the sampling relationship between z and D. Later expressions involving z are evaluated with respect to that sampling convention.

Tracing a Sampled Variable

Suppose D is a distribution over the set Z, and z ∼ D. What does the notation tell you about a later expression f(z)?

Identify D: D is the distribution over the set Z.

Interpret z ∼ D: z is treated as a value sampled according to D rather than as a permanently specified value.

Read f(z): The function is evaluated using the sampled variable z, so its interpretation follows the sampling convention supplied by D.

The full reading is: sample z according to D, then evaluate f using that sampled z.

Expected Values and Events

The expression E z ∼D [f(z)] denotes the expected value of the real-valued function f(z) when z is sampled according to D. The subscript identifies the distribution governing the variable inside the brackets. If the sampling dependence is already clear from context, the shorter notation E[f] may be used; it refers to the same quantity rather than introducing a different one.

Probability notation asks a different question. If f maps Z to true or false, then P z ∼D [f(z)] denotes the probability that the condition f(z) is true when z is sampled according to D. In set notation, this is the probability assigned by D to the collection of values for which f(z) is true.

summarizesmeasuresE z ∼D [f(z)]real-valued functionP z ∼D [f(z)]Boolean conditionexpected valuesummary of f(z)event probabilitycondition is true
What is the difference between summarizing a numerical quantity over samples and measuring how often a Boolean condition is true?

Choosing the Right Notation

You want to describe the numerical result of applying f to a value sampled from D, and separately you want to describe how often a true-or-false condition holds for a value sampled from D. Which notation matches each task?

Numerical quantity: Use E z ∼D [f(z)] when f is a real-valued function and the goal is its expected value under D.

Boolean condition: Use P z ∼D [f(z)] when f maps values to true or false and the goal is the probability that the condition is true.

Compare the roles: The distribution and sampled variable can appear in both expressions, but the outer symbol identifies what is being summarized: an expected value or an event probability.

E indicates an expected value of a real-valued function; P indicates the probability that a Boolean condition is true.

Repeated Independent Sampling

A single sample from D is one point in the set Z. If m points are sampled from D, they form a tuple written as (z₁, …, zₘ). The notation D^m refers to the probability over Z^m induced by this sampling process. Each component zᵢ is sampled from D independently of the other points.

The superscript in D^m does not name a new individual sample variable. It signals a distribution over the larger space of m-component sample tuples. D governs individual samples from Z; D^m governs the resulting tuples in Z^m.

independent drawindependent drawindependent drawcomponentcomponentcomponentinducesDsource distributionz₁sample from Dz₂sample from D(z₁, …, zₘ)point in ZᵐDᵐdistribution over tupleszₘsample from D
How does one distribution D become a joint distribution over m independently sampled points, and how does the data move into the sample tuple?

Interpreting Dᵐ

Suppose m points are sampled independently from D. What object does Dᵐ describe?

List the individual samples: The samples are written as z₁ through zₘ, with each zᵢ sampled from D.

Collect them: Together, the samples form the m-component tuple (z₁, …, zₘ), which is an element of Zᵐ.

Interpret the superscript: Dᵐ describes the probability over those tuples induced by the independent sampling process.

Dᵐ is a distribution over m-component tuples in Zᵐ, not another single point sampled from Z.

Notation Mistakes

  • Treating z ∼ D as if it assigned one permanently fixed value to z.

    The notation describes z as a value obtained by sampling according to D; it does not identify one permanent value.

    Fix: Read it as “sample z according to D,” then interpret later expressions using that sampling context.

  • Using expected-value notation for a Boolean event without noticing the distinction.

    P describes the probability that a Boolean condition is true, while E describes the expected value of a real-valued function.

    Fix: Check whether the expression evaluates a real-valued function or a true-or-false condition.

  • Reading Dᵐ as a single new sample from Z.

    Dᵐ refers to a probability over m-component tuples in Zᵐ formed by independent sampling.

    Fix: Expand the notation into z₁, …, zₘ and identify the tuple (z₁, …, zₘ).

  • Ignoring the distribution named in the expectation subscript.

    The subscript identifies the distribution governing the variable inside the brackets.

    Fix: Include the sampling distribution in your reading: the expected value of f(z) when z is sampled according to D.

Notation Practice

MEDIUM

For each statement, explain what object is being described: D, z ∼ D, E z ∼D [f(z)], P z ∼D [f(z)], and Dᵐ. In your explanation, identify the set involved, the role of sampling, and whether the notation summarizes a numerical function, a Boolean event, or a tuple of samples.

Hints
  • Start with D and identify the set over which it is a distribution.
  • For z ∼ D, say what relationship connects z and D.
  • For E and P, inspect whether the expression concerns a real-valued function or a true-or-false condition.
  • For Dᵐ, expand the object into the samples z₁ through zₘ and the tuple they form.
  1. The central reading pattern is distribution, sample, then evaluation. D is a distribution over a set. z ∼ D says that z is sampled according to D. E z ∼D [f(z)] summarizes the expected value of a real-valued function under that sampling. P z ∼D [f(z)] measures how often a Boolean condition is true. Dᵐ extends the same idea to m independently sampled points and the tuples they form.

Key Takeaways

  • D names a probability distribution over a set, while z ∼ D says that z is sampled according to D.
  • Expected-value notation concerns a real-valued function evaluated on a sampled variable.
  • Probability-of-an-event notation concerns whether a Boolean condition is true for a sampled variable.
  • Dᵐ describes the probability over m-component tuples in Zᵐ produced by independent samples from D.
  • A dependable reading order is to identify D, interpret the sampled variable, and then determine what the outer notation summarizes.