Concepts / Linear Regression

Linear Regression

Regression loss functions measure more than whether a prediction is wrong: they quantify the discrepancy between prediction and target.

  • Programming

Why Wrong Is Not Enough

Imagine estimating a baby's weight from explanatory information such as her age and weight at birth. A prediction of 3.00001 kg and a prediction of 4 kg are both different from an actual value of 3 kg, but they are not equally bad. Regression therefore needs more than a yes-or-no judgment about correctness. It needs a loss function that converts the discrepancy between a prediction and the actual value into a numerical penalty.

A loss function evaluates one prediction against its target. It measures how large the discrepancy is and applies a rule for turning that discrepancy into a penalty. Choosing the loss function is part of choosing the model: it determines what the learning procedure treats as a small error and what it treats as a large error.

Two Ways to Penalize Error

comparecomparesquaretake magnitudePrediction4Discrepancy1Squared loss1Actual value3Absolute value loss1
How do squared loss and absolute value loss assign different penalties to the same prediction error?

Squared loss starts with the prediction error and squares it. Absolute value loss starts with the same error and takes its absolute value. Both begin from the same discrepancy, but they apply different penalty rules. The difference matters because the resulting numerical penalties can be different even when the prediction and actual value are unchanged.

Comparing two penalties

The actual value is 3 kg and the prediction is 4 kg. Calculate the squared loss and the absolute value loss.

Find the discrepancy: The prediction differs from the actual value by 4 minus 3, which is 1 kg.

Apply squared loss: Square the discrepancy: 1 multiplied by 1 gives a squared loss of 1.

Apply absolute value loss: Take the magnitude of the discrepancy: the absolute value of 1 is 1.

For this one-unit error, squared loss is 1 and absolute value loss is 1.

A very small error and a larger error

Compare predictions of 3.00001 kg and 4 kg when the actual value is 3 kg.

Prediction 3.00001: The absolute discrepancy is 0.00001. Its squared loss is 0.0000000001.

Prediction 4: The absolute discrepancy is 1. Its squared loss is 1.

Compare the penalties: Both predictions differ from the target, but the numerical penalties show that their discrepancies are not equally large.

Absolute value loss gives 0.00001 and 1; squared loss gives 0.0000000001 and 1.

From Individual Loss to Empirical Risk

individual lossindividual lossindividual losssquared-loss averageExample 1loss 1Average lossessum divided by countMean Squared Error2Example 2loss 4Example 3loss 1
How do the losses from many individual examples combine into the empirical risk?

A loss function evaluates one example at a time. When squared loss is used across a data set, the resulting empirical risk function is called Mean Squared Error. The individual squared-loss expression is the penalty for one prediction; Mean Squared Error is the data-set-level empirical risk associated with using squared loss.

Computing a data-set-level risk

Suppose three examples have squared losses of 1, 4, and 1. Find the Mean Squared Error.

Add the individual losses: The total is 1 plus 4 plus 1, which equals 6.

Average across examples: There are three examples, so divide 6 by 3.

The Mean Squared Error is 2.

Absolute value loss also has an empirical risk minimization rule, and that rule can be implemented using linear programming. The central distinction remains the same: first define how one example is penalized, then use the chosen loss across the data set when evaluating a predictor.

Inputs, Outcomes, and Predictors

apply hproduce h(x)Input xin X, a subset of R^dLearned predictor hlinear functionPredicted outcomereal number in Y
How does an input from the domain of explanatory variables move through a learned predictor to produce a real-valued label?

Linear regression is a statistical tool for modeling relationships between explanatory variables and a real-valued outcome. The explanatory variables provide the information used as input. The outcome is the quantity the model aims to estimate.

When viewed as a learning problem, the input domain X is a subset of R^d for some d. The label set Y is the set of real numbers. The learning goal is to learn a linear function h that maps an input from R^d to a real number and best approximates the relationship between the explanatory variables and the outcome.

The direction of the problem matters. Explanatory variables are the inputs; the real-valued outcome is the target. A learned predictor uses the input to produce an estimated outcome. Linear regression is therefore not only a way to describe a relationship in data. It is also a way to learn a predictor that approximates that relationship.

The Linear Hypothesis Class

combine withcombine withcombine withCoefficientsweights for input variablesPredicted outcomeh(x)Input variablesx in R^dInterceptconstant term
What do the coefficients and intercept of a linear function control, and how do they determine the predicted outcome?

A hypothesis class is the collection of candidate functions that the learning problem is allowed to consider. In linear regression, the hypothesis class is restricted to linear functions. Learning therefore means selecting or learning a suitable linear function h for approximating the relationship in the data.

A linear function combines input variables with coefficients and an intercept to produce a real-valued prediction. The coefficients determine how the input variables participate in the prediction, while the intercept is the constant part of the function. The important point for the learning setup is that the predictor must come from the permitted class of linear functions.

For the baby-weight problem, age and weight at birth are explanatory variables, and later weight is the real-valued outcome. A learned linear predictor uses the explanatory input to produce an estimated weight. The estimate is judged through a regression loss rather than through a binary correct-or-incorrect test.

Making the Class Finite

discretizeform combinationsContinuousparametersmany possible valuesDiscretized valuesfinite permitted choicesFinite hypothesisclassfinite candidate functions
How does discretizing continuous parameter values turn an effectively infinite set of linear hypotheses into a finite set for sample-complexity analysis?

Linear regression is not a binary prediction task, so its sample complexity cannot be analyzed using the VC-dimension in the same way described for binary classification. One possible route is the discretization trick. Continuous parameter values can be replaced by a finite collection of permitted values. The resulting collection of candidate linear functions is finite, which makes finite-class sample-complexity analysis possible.

A simple discretization idea

Suppose a coefficient is allowed to take any value in a continuous range. Describe what happens if the learning analysis permits only a finite list of coefficient values.

Start with continuous choices: Before discretization, the coefficient can vary over a continuous set of values, so there are many possible linear functions.

Restrict the choices: Replace the continuous range with a finite collection of allowed values.

Form candidate functions: Using finite choices for the parameters produces a finite collection of candidate linear hypotheses.

Discretization changes the analysis from an effectively infinite collection of parameter choices to a finite hypothesis class.

Common Mistakes

  • Treating every incorrect regression prediction as equally bad.

    Regression errors have different sizes, and a loss function is used to quantify those differences.

    Fix: Compute a loss that converts the discrepancy into a numerical penalty.

  • Confusing squared loss on one example with Mean Squared Error.

    Squared loss evaluates one example, while Mean Squared Error is the empirical risk associated with applying squared loss across a data set.

    Fix: Use individual squared loss for one example and Mean Squared Error for the data-set-level average.

  • Reversing the roles of explanatory variables and outcome.

    Explanatory variables provide the input, while the real-valued outcome is what the predictor approximates.

    Fix: Identify the information used for prediction as the input and the estimated quantity as the outcome.

  • Describing linear regression as allowing every possible function.

    The hypothesis class in linear regression is restricted to linear functions.

    Fix: Describe learning as selecting a suitable member of the linear-function hypothesis class.

  • Treating discretization as the definition of linear regression.

    Discretization is one route for sample-complexity analysis, while the modeling task is to learn a linear function for a real-valued outcome.

    Fix: Keep the regression definition and the finite-class analysis technique conceptually separate.

Check Your Understanding

EASY

An actual outcome is 10. A predictor produces 7. Calculate the discrepancy, the squared loss, and the absolute value loss. Then explain whether the result is a loss on one example or an empirical risk across a data set.

Hints
  • Compare the prediction with the actual value first.
  • Squared loss squares the discrepancy.
  • Absolute value loss takes the magnitude of the discrepancy.
  • Empirical risk requires combining losses from multiple examples.
MEDIUM

A learning problem uses inputs from a subset of R^d and real-number labels. Explain why it fits the linear-regression setup, identify the hypothesis class, and describe what the learned predictor is trying to approximate.

Hints
  • The input domain is a subset of R^d.
  • The label set is the set of real numbers.
  • The hypothesis class contains linear functions.
  • The learned predictor approximates the relationship between explanatory variables and the outcome.

Key Takeaways

  1. Regression needs a loss function because prediction errors can differ in size, not merely in whether they are correct or incorrect.
  2. Squared loss squares the discrepancy, while absolute value loss takes its magnitude.
  3. Individual loss evaluates one example; Mean Squared Error is the empirical risk associated with squared loss across a data set.
  4. Linear regression models a relationship between explanatory variables and a real-valued outcome using a hypothesis class of linear functions.
  5. The input domain is a subset of R^d, the label set is the real numbers, and the learned predictor aims to approximate the relationship between input and outcome.
  6. Discretization can replace continuous parameter choices with finite choices, creating a finite hypothesis class for sample-complexity analysis.

Key Takeaways

  • A regression loss measures the size of the discrepancy between a prediction and a real-valued target.
  • Squared loss and absolute value loss use different penalty rules for the same discrepancy.
  • Mean Squared Error is a data-set-level empirical risk, not the loss on one example.
  • Linear regression learns a linear predictor from explanatory variables to a real-valued outcome.
  • Discretization can make the linear-regression hypothesis class finite for sample-complexity analysis.