Concepts / Risk Functions

Risk Functions

The learning objective is to minimize the risk function L_D.

  • Programming

From One Example to a Learning Objective

Learning has a broad objective: minimize the risk function L_D. This objective concerns the overall learning problem, while SGD obtains information about the direction of improvement from one freshly sampled example at a time. The central idea is not that every sampled gradient exactly equals the gradient of L_D. Instead, the expected value of the sampled gradient matches the gradient of the risk function.

evaluateconnect through expectationminimizeParameters w(t)Sampled lossesℓ(w, z)Risk L_DMinimum risk
How does minimizing the risk function connect the current parameters, individual losses, and the overall learning goal?

The Sampled Gradient

At iteration t, SGD uses one fresh sample, written as z. From that sample, it forms the sampled loss ℓ(w, z). The random vector v_t is the gradient of this sampled loss with respect to w, evaluated at the current parameter point w(t). In compact form, v_t is the gradient of ℓ(w, z) with respect to w at w(t).

use zdifferentiateevaluate at w(t)evaluation pointFresh sample zSampled loss ℓ(w, z)Gradient with respectto wRandom vector v_tCurrent point w(t)
How does one sampled training example become the random vector v_t used by SGD?

Tracking the Objects in One SGD Step

Suppose one iteration samples an example z and the current parameter point is w(t). Identify the function, variable, point, and resulting vector in the sampled-gradient construction.

Function: The function being differentiated is the sampled loss ℓ(w, z), not the entire risk function L_D.

Variable: The differentiation variable is w, because the gradient describes how the sampled loss changes as the parameter variable changes.

Point: The gradient is evaluated at the current parameter point w(t).

Result: The resulting gradient vector is called v_t. It is random because it depends on the freshly sampled example z.

v_t is the gradient of ℓ(w, z) with respect to w, evaluated at w(t).

with respect toacts onevaluated atproduces∇gradientwℓ(w, z)sampled lossw(t)evaluation pointv_tresult
Which function is differentiated, which variable is changed, and at what parameter point is the gradient evaluated?

Why One Sample Is Enough to Guide SGD

A single sampled gradient is not required to be the exact gradient of the risk function on that iteration. It is an estimate based only on the sampled loss. Its relevance comes from averaging over the randomness introduced by sampling z: the expected value of v_t matches the gradient of the risk function at the current point. This is the unbiasedness relationship.

contributes to expectationcontributes to expectationequals in expectationv_t from sample zone sampled gradientExpected sampledgradientmatches the risk gradientv_t from another zanother sampled gradientGradient of L_Dat w(t)
How do sampled gradients relate to the true gradient of the risk function, and why does their expected direction match it?
ObjectRoleScope
L_DLearning objective to minimizeRisk function
ℓ(w, z)Loss used for the sampled exampleOne sampled example
v_tRandom gradient vector constructed by SGDCurrent point and one fresh sample
Gradient of L_DDirection represented by the risk objectiveRisk function

The SGD Iteration

  1. Start from the current parameter point w(t).
  2. Sample one fresh example z.
  3. Form the sampled loss ℓ(w, z).
  4. Take its gradient with respect to w.
  5. Evaluate that gradient at w(t) to obtain v_t.
  6. Use v_t as an estimate of the direction represented by the gradient of the risk function.
samplesdefinesdifferentiate at w(t)informs directionSGDFresh example zSampled loss ℓ(w, z)Random vector v_tParameter point w(t)
What happens in sequence when SGD samples an example, computes its gradient, and uses it to address risk minimization?

Mistakes About Risk and Sampling

  • Treating the sampled loss as the risk function.

    The sampled loss is local to one example, while L_D is the risk function that learning seeks to minimize.

    Fix: Keep the target and the sample-level quantity separate: L_D is the objective, and ℓ(w, z) supplies the information used to construct v_t.

  • Assuming v_t must equal the risk gradient for every sample.

    The sampled gradient is an unbiased estimate; its expected value matches the risk gradient, but an individual sample need not produce the exact risk gradient.

    Fix: Interpret equality as an expected-value relationship over the sampling of z.

  • Differentiating with respect to the sampled example.

    The sampled vector is the gradient of ℓ(w, z) with respect to the parameter variable w.

    Fix: Identify w as the variable being changed and w(t) as the point where the gradient is evaluated.

Check Your Understanding

MEDIUM

A learner says: “SGD minimizes the loss ℓ(w, z) for whichever example it sampled, so v_t is the exact gradient of L_D.” Correct this statement using the roles of L_D, ℓ(w, z), v_t, and expectation.

Hints
  • Start by identifying which quantity is the overall learning objective.
  • Then identify which quantity is constructed from one fresh sample.
  • Finally, explain the relationship between one sampled gradient and the gradient of the risk function.

What do you think happens?

If the sampled example changes while the current point w(t) stays fixed, does the resulting vector v_t necessarily stay the same?

  • Yes, because w(t) determines the vector completely.
  • No, because v_t is constructed from the sampled loss for the fresh example.
  • Yes, because every sampled gradient equals the risk gradient.
Reveal answer

Answer: No, because v_t is constructed from the sampled loss for the fresh example.

The vector v_t depends on the sampled example through ℓ(w, z). Its expected value over the sampling of z matches the gradient of the risk function, but individual sampled vectors need not be identical.

Key Takeaways

  1. Learning aims to minimize the risk function L_D.
  2. At each iteration, SGD uses one freshly sampled example z to form the sampled loss ℓ(w, z).
  3. The vector v_t is the gradient of the sampled loss with respect to w, evaluated at the current point w(t).
  4. An individual sampled gradient need not equal the exact gradient of L_D.
  5. The expected value of the sampled vector matches the gradient of the risk function, which connects one-example information to the overall learning objective.

Key Takeaways

  • Risk minimization defines the learning objective through L_D.
  • SGD constructs v_t from the gradient of the sampled loss ℓ(w, z) at w(t).
  • The function, variable, and evaluation point must be distinguished clearly.
  • The sampled gradient is an unbiased estimate in expectation, not necessarily an exact risk gradient for one sample.