Concepts / Mean Squared Value Error

Mean Squared Value Error

SGD performs an immediate, small weight-vector adjustment for each example.

  • Programming

An Immediate Correction

Imagine a value-prediction system receiving examples one after another. A stochastic gradient descent method does not wait until every example has been processed before changing its parameters. Instead, it makes a small adjustment to its weight vector immediately after each example. The immediate adjustment responds to the error associated with the example just considered.

The word stochastic refers here to using one example for each update.

supplies local error informationstarts fromproducesExampleone observed caseθtcurrent weight vectorSmall adjustmentnegative gradient scaled byαθt+1updated weight vector
What changes in the weight vector immediately after processing one training example?

Reading the Update

θt+1 = θt − α∇f(θt)

In this update, θt is the current weight vector and θt+1 is the vector after the update. The symbol α is a positive step-size parameter. The term ∇f(θt) is the gradient of the chosen scalar expression with respect to the components of the weight vector.

PartMeaning
θtThe current weight vector
θt+1The weight vector after the update
αA positive parameter controlling the scale of the step
∇f(θt)The gradient, with one partial derivative for each weight component
−∇f(θt)The direction used for decreasing the error

The roles of the terms in a gradient-descent update.

The gradient is a derivative vector rather than a single number. Each component describes how the error changes when the corresponding weight changes. Together, the components identify the direction of greatest increase in the error. The subtraction in the update is therefore essential: moving in the opposite direction points toward the most rapid decrease indicated by the gradient.

evaluate error changereverse directiontake a stepWeight vectorcurrent positionGradientsteepest error increaseNegative gradientmost rapid decreasedirectionDescent updatescaled by α
How does the gradient identify the direction where error increases most steeply, and why does descent move oppositely?

Tracing a Single Step

A Two-Component Weight Vector

Suppose the current weight vector is θt = (4, 1), the gradient for the example is ∇f(θt) = (2, −1), and the step size is α = 0.1. Apply the update θt+1 = θt − α∇f(θt).

Scale the gradient: Multiplying the gradient by the step size gives α∇f(θt) = 0.1(2, −1) = (0.2, −0.1).

Reverse the direction: The update subtracts this scaled gradient, so the adjustment is (−0.2, 0.1).

Apply the adjustment: Add the adjustment to the current vector: (4, 1) + (−0.2, 0.1).

The updated weight vector is θt+1 = (3.8, 1.1).

This example shows two separate roles in the update. The gradient determines the direction of the change, while α determines how far the weight vector moves in that direction. A larger positive step size would scale the same gradient more strongly; a smaller positive step size would produce a smaller adjustment.

From Local Steps to MSVE

A single stochastic update focuses on one example's error. It does not directly calculate an adjustment using every example at once. The connection to Mean Squared Value Error comes from repetition: the method processes many examples, making a small correction after each one. These local corrections can collectively reduce an average objective such as MSVE.

processadjust immediatelycontinueadjust immediatelyrepeat across examplessupportCurrent weightsθtExample 1local squared errorSmall updatefirst correctionExample 2local squared errorSmall updatenext correctionMany examplesrepeated local correctionsLower MSVEreduced average squarederror
How do many small updates from individual examples combine to reduce the average squared value error?

The local and broad interpretations should not be confused. Locally, each update tries to reduce the squared error for the example currently being considered. Broadly, repeated updates across examples can reduce the average error measured by MSVE. When the examples follow the same distribution as the states used for the MSVE objective, corrections based on observed examples can support that average-performance goal.

FeatureOrdinary gradient descentStochastic gradient descent
When parameters changeAfter the method has processed every exampleImmediately after each example
Immediate directionBased on the full set of examplesBased on one example
Connection to MSVEUses an update intended to address the overall objective directlyUses repeated local updates that can reduce the average objective

Mistakes About Stochastic Updates

  • Thinking stochastic means that the weight vector changes only occasionally or unpredictably.

    In this context, stochastic refers to using one example for each update, and the adjustment happens immediately after that example.

    Fix: Treat each processed example as an opportunity for an immediate small weight-vector adjustment.

  • Moving in the direction of the gradient when trying to reduce error.

    The gradient points toward the direction of greatest error increase.

    Fix: Use the negative gradient for a descent step, as in θt+1 = θt − α∇f(θt).

  • Treating the gradient as a single number.

    The gradient is a derivative vector with one partial derivative for each component of the weight vector.

    Fix: Read each gradient component as information about the corresponding weight component.

  • Assuming that one example directly minimizes MSVE.

    A single update focuses on one local example error, while MSVE is an average objective supported by repeated corrections across examples.

    Fix: Separate the local purpose of one update from the broader effect of many updates.

Check Your Reasoning

MEDIUM

A stochastic method has current weight vector θt, processes one example, and computes a gradient for that example. Explain what information determines the direction of the update, what determines its scale, and why repeating this process over many examples can support reducing MSVE.

Hints
  • The direction comes from the sign-reversed gradient.
  • The step-size parameter controls the scale.
  • Connect one example's local squared error with repeated updates and an average objective.

What do you think happens?

If the same gradient is used twice but the positive step size is made smaller, what changes?

  • The update direction reverses
  • The update moves a shorter distance in the same descent direction
  • The gradient becomes a scalar
  • The method waits for the full dataset
Reveal answer

Answer: The update moves a shorter distance in the same descent direction.

The negative gradient determines the descent direction, while the positive step-size parameter controls the scale of the update.

Key Takeaways

  1. Stochastic gradient descent updates the weight vector immediately after each example rather than waiting for all examples.
  2. The gradient is a vector describing how the error changes with respect to the weight components.
  3. The gradient points toward the steepest error increase, so descent uses its negative.
  4. The positive step-size parameter controls how large each update is.
  5. Each update addresses one example locally, while many updates across suitable examples can reduce an average objective such as MSVE.

Key Takeaways

  • SGD makes a small weight-vector adjustment after every example.
  • The negative gradient supplies the descent direction, and the step size controls the distance moved.
  • A single stochastic update is local to one example rather than a direct update over the entire dataset.
  • Repeated local corrections can support reducing the average squared value error measured by MSVE.