Risk Functions
The learning objective is to minimize the risk function L_D.
From One Example to a Learning Objective
Learning has a broad objective: minimize the risk function L_D. This objective concerns the overall learning problem, while SGD obtains information about the direction of improvement from one freshly sampled example at a time. The central idea is not that every sampled gradient exactly equals the gradient of L_D. Instead, the expected value of the sampled gradient matches the gradient of the risk function.
The Sampled Gradient
At iteration t, SGD uses one fresh sample, written as z. From that sample, it forms the sampled loss ℓ(w, z). The random vector v_t is the gradient of this sampled loss with respect to w, evaluated at the current parameter point w(t). In compact form, v_t is the gradient of ℓ(w, z) with respect to w at w(t).
Tracking the Objects in One SGD Step
Suppose one iteration samples an example z and the current parameter point is w(t). Identify the function, variable, point, and resulting vector in the sampled-gradient construction.
Function: The function being differentiated is the sampled loss ℓ(w, z), not the entire risk function L_D.
Variable: The differentiation variable is w, because the gradient describes how the sampled loss changes as the parameter variable changes.
Point: The gradient is evaluated at the current parameter point w(t).
Result: The resulting gradient vector is called v_t. It is random because it depends on the freshly sampled example z.
v_t is the gradient of ℓ(w, z) with respect to w, evaluated at w(t).
Why One Sample Is Enough to Guide SGD
A single sampled gradient is not required to be the exact gradient of the risk function on that iteration. It is an estimate based only on the sampled loss. Its relevance comes from averaging over the randomness introduced by sampling z: the expected value of v_t matches the gradient of the risk function at the current point. This is the unbiasedness relationship.
| Object | Role | Scope |
|---|---|---|
| L_D | Learning objective to minimize | Risk function |
| ℓ(w, z) | Loss used for the sampled example | One sampled example |
| v_t | Random gradient vector constructed by SGD | Current point and one fresh sample |
| Gradient of L_D | Direction represented by the risk objective | Risk function |
The SGD Iteration
- Start from the current parameter point w(t).
- Sample one fresh example z.
- Form the sampled loss ℓ(w, z).
- Take its gradient with respect to w.
- Evaluate that gradient at w(t) to obtain v_t.
- Use v_t as an estimate of the direction represented by the gradient of the risk function.
Mistakes About Risk and Sampling
Treating the sampled loss as the risk function.
The sampled loss is local to one example, while L_D is the risk function that learning seeks to minimize.
Fix:
Keep the target and the sample-level quantity separate: L_D is the objective, and ℓ(w, z) supplies the information used to construct v_t.Assuming v_t must equal the risk gradient for every sample.
The sampled gradient is an unbiased estimate; its expected value matches the risk gradient, but an individual sample need not produce the exact risk gradient.
Fix:
Interpret equality as an expected-value relationship over the sampling of z.Differentiating with respect to the sampled example.
The sampled vector is the gradient of ℓ(w, z) with respect to the parameter variable w.
Fix:
Identify w as the variable being changed and w(t) as the point where the gradient is evaluated.
Check Your Understanding
A learner says: “SGD minimizes the loss ℓ(w, z) for whichever example it sampled, so v_t is the exact gradient of L_D.” Correct this statement using the roles of L_D, ℓ(w, z), v_t, and expectation.
Hints
- Start by identifying which quantity is the overall learning objective.
- Then identify which quantity is constructed from one fresh sample.
- Finally, explain the relationship between one sampled gradient and the gradient of the risk function.
What do you think happens?
If the sampled example changes while the current point w(t) stays fixed, does the resulting vector v_t necessarily stay the same?
Reveal answer
Answer: No, because v_t is constructed from the sampled loss for the fresh example.
The vector v_t depends on the sampled example through ℓ(w, z). Its expected value over the sampling of z matches the gradient of the risk function, but individual sampled vectors need not be identical.
Key Takeaways
- Learning aims to minimize the risk function L_D.
- At each iteration, SGD uses one freshly sampled example z to form the sampled loss ℓ(w, z).
- The vector v_t is the gradient of the sampled loss with respect to w, evaluated at the current point w(t).
- An individual sampled gradient need not equal the exact gradient of L_D.
- The expected value of the sampled vector matches the gradient of the risk function, which connects one-example information to the overall learning objective.
Key Takeaways
- Risk minimization defines the learning objective through L_D.
- SGD constructs v_t from the gradient of the sampled loss ℓ(w, z) at w(t).
- The function, variable, and evaluation point must be distinguished clearly.
- The sampled gradient is an unbiased estimate in expectation, not necessarily an exact risk gradient for one sample.