Concepts / Mean Squared Error

Mean Squared Error

Regression loss functions measure more than whether a prediction is wrong: they quantify the discrepancy between prediction and target.

  • Programming

Why Error Size Matters

A regression model rarely produces only perfectly correct or completely incorrect answers. Suppose the actual value is 3 kg. A prediction of 3.00001 kg and a prediction of 4 kg are both different from the actual value, but they are not equally bad. Regression therefore needs a rule that converts the discrepancy between a prediction and the actual value into a numerical penalty.

comparecomparecompareActual value 3 kgPrediction 3 kgno discrepancyPrediction 3.00001 kgsmall discrepancyPrediction 4 kglarger discrepancy
What happens to a regression loss as the prediction moves from exactly correct to slightly wrong to substantially wrong?

Two Penalty Rules

A loss function evaluates one prediction against its target. The prediction and target provide the same discrepancy in either calculation, but the loss function determines how that discrepancy becomes a penalty. Two common choices for regression are squared loss and absolute value loss.

Loss choiceHow the discrepancy is transformedWhat changes when the discrepancy changes
Squared lossSquare the discrepancyThe penalty follows the squared-loss rule
Absolute value lossTake the absolute value of the discrepancyThe penalty follows the absolute-value rule

Both choices begin with the prediction-target discrepancy but apply different penalty rules.

apply ruleapply ruleapply ruleapply ruleSquared lossdiscrepancy squaredSmall discrepancyAbsolute value lossabsolute discrepancyLarge discrepancy
How do squared loss and absolute value loss respond differently as the prediction moves farther from the actual value?

Applying Loss to One Example

One Prediction, Two Losses

A model predicts 5 for a target whose actual value is 3. Apply squared loss and absolute value loss.

Find the discrepancy: Subtract the actual value from the prediction: 5 - 3 = 2.

Apply squared loss: Square the discrepancy: 2 squared = 4.

Apply absolute value loss: Take the absolute value of the discrepancy: absolute value of 2 = 2.

For this generated example, squared loss is 4 and absolute value loss is 2. The two values differ because the same discrepancy was passed through two different penalty rules.

combine with targetcombine with predictionsquareabsolute valuePrediction 5Discrepancy 25 minus 3Squared loss 42 squaredTarget 3Absolute loss 2absolute value of 2
How is one numerical prediction-target discrepancy transformed into squared loss or absolute value loss?

What do you think happens?

A prediction is 7 and the target is 3. Which loss is larger for this example?

  • Squared loss
  • Absolute value loss
  • They are always equal
Reveal answer

Answer: Squared loss

The discrepancy is 4. Squared loss applies the squaring rule to that discrepancy, while absolute value loss applies the absolute-value rule. The penalty values therefore differ.

From One Loss to Mean Squared Error

A loss function evaluates one example at a time. When squared loss is used across a data set, the resulting empirical risk function is called the Mean Squared Error. The individual squared-loss expression describes one prediction's penalty; Mean Squared Error names the empirical risk associated with using squared loss across the sample.

Individual Loss Versus Empirical Risk

Consider two generated prediction-target pairs. The first pair has prediction 5 and target 3. The second pair has prediction 4 and target 5. Compare the squared loss on the first example with the empirical risk obtained by averaging the two squared losses.

Evaluate the first example: Its discrepancy is 5 - 3 = 2, so its squared loss is 4.

Evaluate the second example: Its discrepancy is 4 - 5 = -1, so its squared loss is 1 when the discrepancy is squared.

Combine the sample losses: The empirical risk associated with squared loss averages the squared losses across the data set: (4 + 1) divided by 2 = 2.5.

Compare the quantities: The first example's squared loss is 4. The sample-level empirical risk is 2.5. They are different quantities because one describes one example and the other averages losses across the sample.

Individual squared loss: 4. Mean Squared Error as the empirical risk for this two-example sample: 2.5.

average squared lossesone contributionOne examplesquared loss 4Mean Squared Errorempirical risk 2.5Data settwo examples
How does the loss for one prediction differ from the empirical risk obtained by averaging losses across an entire sample?

Discretization for Analysis

Loss functions also matter when studying how much data is needed for learning guarantees. Linear regression is not a binary prediction task, so its sample complexity cannot be analyzed using the VC-dimension in the way described for binary classification. One possible route is the discretization trick.

The basic idea is to replace a continuous collection of possible linear hypotheses with a finite collection by discretizing possible parameter or prediction values. Instead of allowing every value in a continuous range, the analysis considers selected discrete values. This produces a finite hypothesis class that can be used in a sample-complexity analysis.

replace continuous choicesobtainContinuous linearhypothesesmany possible parametervaluesDiscretizationselect discrete valuesFinite hypothesisclassfinite collection
How does discretizing possible parameter or prediction values change a continuous set of linear hypotheses into a finite collection?

Mistakes to Avoid

  • Treating every wrong regression prediction as equally bad.

    Regression loss is intended to quantify the discrepancy, not merely record whether the prediction is different.

    Fix: Use a loss function that converts the size of the discrepancy into a numerical penalty.

  • Confusing squared loss with absolute value loss.

    The two choices begin with the same discrepancy but apply different penalty rules.

    Fix: Identify the selected loss function before transforming the discrepancy.

  • Calling the loss on one example Mean Squared Error.

    Mean Squared Error is the empirical risk function associated with squared loss across a data set.

    Fix: Use squared loss for one example and Mean Squared Error for the empirical-risk calculation across the sample.

  • Assuming binary-classification analysis applies directly to linear regression.

    Linear regression is not a binary prediction task.

    Fix: Recognize discretization as one possible route to making the linear-regression hypothesis class finite for analysis.

Check Your Understanding

MEDIUM

A model makes two predictions. For the first example, the prediction is 6 and the target is 4. For the second, the prediction is 3 and the target is 4. Calculate the squared loss for each example, then calculate the empirical risk associated with squared loss by averaging the two values. Finally, state whether your individual loss for the first example is the same quantity as the empirical risk.

Hints
  • Compute each prediction-target discrepancy first.
  • Apply the squaring rule to each discrepancy.
  • Average the two squared losses for the empirical-risk value.
  • Keep the one-example loss separate from the sample-level average.
MEDIUM

In your own words, explain why replacing continuous parameter or prediction values with discrete choices can help make a linear-regression hypothesis class finite for sample-complexity analysis.

Hints
  • Start with the difference between continuous choices and a finite collection.
  • Connect the finite collection to the purpose of sample-complexity analysis.

Key Takeaways

  1. Regression needs a loss function because predictions can be slightly or substantially wrong, and those discrepancies should receive different numerical penalties.
  2. Squared loss and absolute value loss begin with the same prediction-target discrepancy but transform it using different rules.
  3. Mean Squared Error is the empirical risk associated with squared loss across a data set, not the loss on only one example.
  4. Discretization can replace a continuous collection of possible linear hypotheses with a finite hypothesis class for sample-complexity analysis.

Key Takeaways

  • A regression loss measures the size of the discrepancy between a prediction and its target.
  • Squared loss squares the discrepancy, while absolute value loss takes its absolute value.
  • The loss on one example and the empirical risk across a data set are different quantities.
  • Mean Squared Error is the empirical risk function associated with squared loss.
  • Discretization can make a linear-regression hypothesis class finite for sample-complexity analysis.