Concepts / Differentiable Optimization

Differentiable Optimization

The learning objective is to minimize the risk function L_D.

  • Programming

From Learning Goal to Update

Learning begins with a broad objective: minimize the risk function L_D. Stochastic Gradient Descent, or SGD, addresses this objective by using one fresh example at each iteration. Instead of directly constructing the full gradient of the risk from all examples, SGD constructs a random gradient vector from the sampled example and uses that vector as an estimate of the direction represented by the risk gradient at the current parameter value.

Keep the target and the estimate separate: L_D is the risk function that learning seeks to minimize, while v_t is the random vector constructed from one sampled example.

One Sample, One Stochastic Vector

At iteration t, SGD considers the current point w(t) and draws one fresh example, written as z. The sampled example determines a sampled loss, written as ℓ(w, z). The gradient is taken with respect to w, and the resulting gradient is evaluated at the current point w(t). That evaluated gradient is the random vector v_t.

formsdifferentiateevaluate atproducesSample zone fresh exampleℓ(w, z)sampled lossGradientwith respect to ww(t)current pointv_tevaluated sampled gradient
How does one sampled training example provide the ingredients for the stochastic gradient vector?

The random vector v_t is the gradient of the sampled loss ℓ(w, z), evaluated at the current point w(t).

Reading the Gradient Precisely

differentiateevaluate atgivesℓ(w, z)sampled losswdifferentiate with respecttow(t)evaluation pointv_trandom gradient vector
Which function is differentiated, with respect to which variable, and at what point is the resulting gradient evaluated?

Tracing one SGD construction

Describe what happens when SGD uses one freshly sampled example z at the current point w(t).

Identify the sample: The iteration uses one fresh example, denoted by z.

Identify the function: That example defines the sampled loss ℓ(w, z). The gradient is taken from this sampled loss rather than directly from the entire risk function.

Identify the variable: The differentiation is with respect to w, the variable representing the current model parameters.

Identify the point: The resulting gradient is evaluated at the current parameter value w(t).

Name the result: The evaluated sampled gradient is the random vector v_t.

One fresh example supplies z, z determines ℓ(w, z), the loss is differentiated with respect to w, and the gradient is evaluated at w(t) to construct v_t.

Why Sampling Still Represents Risk

A single sampled gradient does not necessarily equal the exact gradient of the risk function. Its relevance comes from considering the expected value of the random vector over the sampling of z. The expected sampled gradient matches the gradient of the risk function. Therefore, although an individual v_t can differ from the full risk gradient, the sampling construction is unbiased with respect to that risk gradient.

contributescontributescontributesmatchesv_t from z_aone sampled gradientExpected v_tover sampling of zGradient of L_Drisk gradientv_t from z_banother sampled gradientv_t from z_canother sampled gradient
How do gradients from individual sampled examples average out to the gradient of the full risk function?
ObjectRole
L_DThe risk function that learning seeks to minimize
ℓ(w, z)The loss associated with one sampled example z
v_tThe random gradient vector obtained from the sampled loss at w(t)
Expected v_tAn unbiased estimate of the gradient of the risk function

Unbiased does not mean identical on every iteration. It means that the expected value of the sampled vector matches the gradient of the risk function.

Repeated Parameter Updates

SGD repeats the same connection at successive iterations. At each iteration, the current point w(t) is paired with one freshly sampled example. That example produces a sampled loss and, from it, the random vector v_t. The vector supplies an estimate of the direction represented by the gradient of the risk at the current point. The process therefore connects the broad objective of minimizing L_D with information obtained one fresh example at a time.

sampleconstructinform next pointsampleconstructinform next pointw(0)current pointFresh ziteration 1v_1sampled gradientw(1)next pointFresh ziteration 2v_2sampled gradientw(2)later point
What happens to the current parameter point as each sampled gradient supplies information for the next update?

The important mechanism is not that every individual sampled vector is the exact risk gradient. The mechanism is that each vector is constructed at the current point from a fresh example, and the expected value of this random construction corresponds to the risk gradient.

Common Reasoning Mistakes

  • Treating v_t as the exact gradient of L_D for every sampled example.

    The sampled gradient is an unbiased estimate, not necessarily the exact risk gradient on one iteration.

    Fix: Distinguish an individual sampled vector from its expected value over the sampling of z.

  • Differentiating the entire risk function when describing the construction of v_t.

    The construction is local to the sampled example.

    Fix: Start with the sampled loss ℓ(w, z), differentiate with respect to w, and evaluate at w(t).

  • Leaving out the evaluation point.

    The random vector is the gradient evaluated at the current point, not an unspecified gradient.

    Fix: Name all three parts: the function ℓ(w, z), the variable w, and the point w(t).

  • Confusing the learning objective with the sampling procedure.

    The objective is to minimize L_D; the sampled example is the source of an estimate used to address that objective.

    Fix: Label L_D as the target and v_t as the estimate.

Practice the Construction

EASY

A fresh example z is sampled while the current parameter point is w(t). Explain, in order, which loss is used, which variable is differentiated, where the gradient is evaluated, and what the resulting vector is called.

Hints
  • The sampled loss is written ℓ(w, z).
  • The gradient is taken with respect to w.
  • The evaluation point is w(t).
  • The resulting random vector is v_t.
MEDIUM

Explain why it is correct to use one sampled gradient even though that vector does not necessarily equal the exact gradient of L_D.

Hints
  • Focus on the expected value of the random vector.
  • State what that expected value matches.
  • Keep the individual sampled vector distinct from the expectation over samples.

Essential Takeaways

  1. Learning seeks to minimize the risk function L_D.
  2. SGD uses one fresh sampled example at each iteration.
  3. The sampled example produces the loss ℓ(w, z), which is differentiated with respect to w and evaluated at w(t) to construct v_t.
  4. An individual v_t need not equal the exact gradient of L_D.
  5. The expected value of the sampled gradient matches the gradient of the risk function, making the sampled construction an unbiased estimate.

Key Takeaways

  • Risk minimization defines the learning objective through the risk function L_D.
  • SGD constructs a random vector v_t from one freshly sampled example.
  • The construction differentiates the sampled loss ℓ(w, z) with respect to w and evaluates the result at w(t).
  • The sampled vector is not necessarily the exact risk gradient for one sample.
  • Its expected value matches the gradient of the risk function, which explains its role in SGD.