Linear Regression
Regression loss functions measure more than whether a prediction is wrong: they quantify the discrepancy between prediction and target.
Why Wrong Is Not Enough
Imagine estimating a baby's weight from explanatory information such as her age and weight at birth. A prediction of 3.00001 kg and a prediction of 4 kg are both different from an actual value of 3 kg, but they are not equally bad. Regression therefore needs more than a yes-or-no judgment about correctness. It needs a loss function that converts the discrepancy between a prediction and the actual value into a numerical penalty.
A loss function evaluates one prediction against its target. It measures how large the discrepancy is and applies a rule for turning that discrepancy into a penalty. Choosing the loss function is part of choosing the model: it determines what the learning procedure treats as a small error and what it treats as a large error.
Two Ways to Penalize Error
Squared loss starts with the prediction error and squares it. Absolute value loss starts with the same error and takes its absolute value. Both begin from the same discrepancy, but they apply different penalty rules. The difference matters because the resulting numerical penalties can be different even when the prediction and actual value are unchanged.
Comparing two penalties
The actual value is 3 kg and the prediction is 4 kg. Calculate the squared loss and the absolute value loss.
Find the discrepancy: The prediction differs from the actual value by 4 minus 3, which is 1 kg.
Apply squared loss: Square the discrepancy: 1 multiplied by 1 gives a squared loss of 1.
Apply absolute value loss: Take the magnitude of the discrepancy: the absolute value of 1 is 1.
For this one-unit error, squared loss is 1 and absolute value loss is 1.
A very small error and a larger error
Compare predictions of 3.00001 kg and 4 kg when the actual value is 3 kg.
Prediction 3.00001: The absolute discrepancy is 0.00001. Its squared loss is 0.0000000001.
Prediction 4: The absolute discrepancy is 1. Its squared loss is 1.
Compare the penalties: Both predictions differ from the target, but the numerical penalties show that their discrepancies are not equally large.
Absolute value loss gives 0.00001 and 1; squared loss gives 0.0000000001 and 1.
From Individual Loss to Empirical Risk
A loss function evaluates one example at a time. When squared loss is used across a data set, the resulting empirical risk function is called Mean Squared Error. The individual squared-loss expression is the penalty for one prediction; Mean Squared Error is the data-set-level empirical risk associated with using squared loss.
Computing a data-set-level risk
Suppose three examples have squared losses of 1, 4, and 1. Find the Mean Squared Error.
Add the individual losses: The total is 1 plus 4 plus 1, which equals 6.
Average across examples: There are three examples, so divide 6 by 3.
The Mean Squared Error is 2.
Absolute value loss also has an empirical risk minimization rule, and that rule can be implemented using linear programming. The central distinction remains the same: first define how one example is penalized, then use the chosen loss across the data set when evaluating a predictor.
Inputs, Outcomes, and Predictors
Linear regression is a statistical tool for modeling relationships between explanatory variables and a real-valued outcome. The explanatory variables provide the information used as input. The outcome is the quantity the model aims to estimate.
When viewed as a learning problem, the input domain X is a subset of R^d for some d. The label set Y is the set of real numbers. The learning goal is to learn a linear function h that maps an input from R^d to a real number and best approximates the relationship between the explanatory variables and the outcome.
The direction of the problem matters. Explanatory variables are the inputs; the real-valued outcome is the target. A learned predictor uses the input to produce an estimated outcome. Linear regression is therefore not only a way to describe a relationship in data. It is also a way to learn a predictor that approximates that relationship.
The Linear Hypothesis Class
A hypothesis class is the collection of candidate functions that the learning problem is allowed to consider. In linear regression, the hypothesis class is restricted to linear functions. Learning therefore means selecting or learning a suitable linear function h for approximating the relationship in the data.
A linear function combines input variables with coefficients and an intercept to produce a real-valued prediction. The coefficients determine how the input variables participate in the prediction, while the intercept is the constant part of the function. The important point for the learning setup is that the predictor must come from the permitted class of linear functions.
For the baby-weight problem, age and weight at birth are explanatory variables, and later weight is the real-valued outcome. A learned linear predictor uses the explanatory input to produce an estimated weight. The estimate is judged through a regression loss rather than through a binary correct-or-incorrect test.
Making the Class Finite
Linear regression is not a binary prediction task, so its sample complexity cannot be analyzed using the VC-dimension in the same way described for binary classification. One possible route is the discretization trick. Continuous parameter values can be replaced by a finite collection of permitted values. The resulting collection of candidate linear functions is finite, which makes finite-class sample-complexity analysis possible.
A simple discretization idea
Suppose a coefficient is allowed to take any value in a continuous range. Describe what happens if the learning analysis permits only a finite list of coefficient values.
Start with continuous choices: Before discretization, the coefficient can vary over a continuous set of values, so there are many possible linear functions.
Restrict the choices: Replace the continuous range with a finite collection of allowed values.
Form candidate functions: Using finite choices for the parameters produces a finite collection of candidate linear hypotheses.
Discretization changes the analysis from an effectively infinite collection of parameter choices to a finite hypothesis class.
Common Mistakes
Treating every incorrect regression prediction as equally bad.
Regression errors have different sizes, and a loss function is used to quantify those differences.
Fix:
Compute a loss that converts the discrepancy into a numerical penalty.Confusing squared loss on one example with Mean Squared Error.
Squared loss evaluates one example, while Mean Squared Error is the empirical risk associated with applying squared loss across a data set.
Fix:
Use individual squared loss for one example and Mean Squared Error for the data-set-level average.Reversing the roles of explanatory variables and outcome.
Explanatory variables provide the input, while the real-valued outcome is what the predictor approximates.
Fix:
Identify the information used for prediction as the input and the estimated quantity as the outcome.Describing linear regression as allowing every possible function.
The hypothesis class in linear regression is restricted to linear functions.
Fix:
Describe learning as selecting a suitable member of the linear-function hypothesis class.Treating discretization as the definition of linear regression.
Discretization is one route for sample-complexity analysis, while the modeling task is to learn a linear function for a real-valued outcome.
Fix:
Keep the regression definition and the finite-class analysis technique conceptually separate.
Check Your Understanding
An actual outcome is 10. A predictor produces 7. Calculate the discrepancy, the squared loss, and the absolute value loss. Then explain whether the result is a loss on one example or an empirical risk across a data set.
Hints
- Compare the prediction with the actual value first.
- Squared loss squares the discrepancy.
- Absolute value loss takes the magnitude of the discrepancy.
- Empirical risk requires combining losses from multiple examples.
A learning problem uses inputs from a subset of R^d and real-number labels. Explain why it fits the linear-regression setup, identify the hypothesis class, and describe what the learned predictor is trying to approximate.
Hints
- The input domain is a subset of R^d.
- The label set is the set of real numbers.
- The hypothesis class contains linear functions.
- The learned predictor approximates the relationship between explanatory variables and the outcome.
Key Takeaways
- Regression needs a loss function because prediction errors can differ in size, not merely in whether they are correct or incorrect.
- Squared loss squares the discrepancy, while absolute value loss takes its magnitude.
- Individual loss evaluates one example; Mean Squared Error is the empirical risk associated with squared loss across a data set.
- Linear regression models a relationship between explanatory variables and a real-valued outcome using a hypothesis class of linear functions.
- The input domain is a subset of R^d, the label set is the real numbers, and the learned predictor aims to approximate the relationship between input and outcome.
- Discretization can replace continuous parameter choices with finite choices, creating a finite hypothesis class for sample-complexity analysis.
Key Takeaways
- A regression loss measures the size of the discrepancy between a prediction and a real-valued target.
- Squared loss and absolute value loss use different penalty rules for the same discrepancy.
- Mean Squared Error is a data-set-level empirical risk, not the loss on one example.
- Linear regression learns a linear predictor from explanatory variables to a real-valued outcome.
- Discretization can make the linear-regression hypothesis class finite for sample-complexity analysis.