Concepts / Squared Loss

Squared Loss

For linear regression with squared loss, least squares solves the empirical risk minimization problem.

  • Programming

From Prediction to Optimization

Linear regression chooses a predictor from a class of possible predictors. The training data determine which predictor is best through empirical risk minimization: choose the predictor whose loss on the training set is smallest. When the loss is squared loss, least squares is the algorithm used to perform this search for the linear regression class.

select fromevaluatecombineminimizeTraining dataLinear predictorsSquared lossesSquared-lossobjectiveLeast squares
How do training examples become one squared-loss objective that least squares minimizes?

The important transition is from many training examples to one objective function. Each example contributes to the squared-loss objective, and least squares searches for the regression weights that minimize that objective.

Residuals Become Loss

A regression predictor produces a prediction for each training instance. The difference between a target and its prediction is a residual error. Under squared loss, the contribution of that error is squared before the contributions are combined into the objective that least squares minimizes.

squarecombineResidual errortarget minus predictionSquared losssquared contributionLoss objectivecombined contributions
How does each prediction error become a squared contribution to the total loss?

Tracing One Training Problem

A linear regression model must be selected from a class of possible predictors, and squared loss is used on the training set. What sequence of ideas turns this selection problem into a least-squares problem?

Evaluate predictors: Use the training data to evaluate the loss produced by possible linear predictors.

Form the objective: Use the squared-loss contributions from the training instances to define the squared-loss objective.

Choose least squares: Least squares carries out the empirical risk minimization search for the linear regression class when the loss is squared loss.

Find the weights: The regression weights are obtained by deriving the gradient condition and rewriting it as a linear system.

The original predictor-selection problem becomes the task of finding regression weights that minimize the squared-loss objective.

Why the Gradient Vanishes

Least squares minimizes the squared-loss objective. To locate an optimal set of weights, the derivation calculates the gradient of that objective and compares it with zero. The zero-gradient condition is not merely a diagnostic checked after training. It is the mathematical condition used to derive the system that determines the weights.

evaluatedifferentiatecompare with zerorewriteWeight vector wSquared-lossobjectiveObjective gradientGradient equals zeroA w = b
How does changing the weight vector lead to the condition used to identify the regression weights?

The System A w = b

After the gradient condition has been formed, the empirical risk minimization problem can be rewritten as the linear system A w = b. In this compact expression, A and b are determined by the problem formulation, while w is the vector of regression weights that must be found.

determinesdeterminescoefficient matrixunknown vectorright-hand sideProblem formulationMatrix Adetermined by formulationVector bdetermined by formulationWeight vector wunknown to findA w = blinear system
How do the problem formulation and the unknown weights meet in the central system A w = b?

Reading the Linear System

Interpret the roles of A, w, and b in A w = b after the gradient condition has been derived.

Identify A: A is determined by the formulation of the regression problem.

Identify b: b is also determined by the formulation of the regression problem.

Identify w: w is the vector of regression weights that must be found.

Solve the system: The solution of A w = b gives the regression weights for the least-squares problem.

A w = b is the central computational form: A and b describe the formed system, and w is the unknown weight vector.

When A Is Invertible

Once the gradient condition has been converted into A w = b, the next question concerns matrix A. If A is invertible, the empirical risk minimization problem has the corresponding direct linear-algebra solution for the weights. The optimization problem has therefore been reduced to solving the system under the invertible-matrix case.

solve directlyanalyze withInvertible ANon-invertible ADirect solutionregression weightsEigenvaluedecompositionlinear-algebra tools
What changes in the solution method when A has an inverse compared with when A is not invertible?

When A Is Not Invertible

A different tool is needed when A is not invertible. This situation can occur when the training instances do not span the entire space of R d. Eigenvalue decomposition provides the needed linear-algebra tools for analyzing this case.

paired withhelps identifyhelps identifyEigenvaluepaired with eigenvectorEigenvectorparameter directionConstraineddirectionsrevealed by decompositionDegenerate directionsrevealed by decomposition
How can eigenvalues and eigenvectors organize the directions that the non-invertible system can constrain or leave degenerate?

The non-invertible case does not discard the original system. The system remains A w = b, and in this setting it still has a solution because b is in the range of A. Eigenvalue decomposition supplies the linear-algebra perspective needed to work with the structure that makes A non-invertible.

Mistakes to Avoid

  • Treating least squares as separate from empirical risk minimization.

    For linear regression with squared loss, least squares is the algorithm used to carry out empirical risk minimization.

    Fix: Start with the training objective: choose the linear predictor whose squared loss on the training set is smallest.

  • Skipping the gradient condition.

    The gradient condition is the mathematical step used to derive the linear system.

    Fix: Calculate the objective gradient, compare it with zero, and then rewrite that condition as A w = b.

  • Assuming A is always invertible.

    The source distinguishes an invertible case from a non-invertible case, which requires eigenvalue-decomposition tools.

    Fix: Identify the status of A before choosing the linear-algebra treatment.

  • Thinking that non-invertibility removes the system.

    The system remains the central form, and in the described setting it still has a solution because b is in the range of A.

    Fix: Keep the system and use eigenvalue decomposition to analyze the non-invertible case.

Practice the Derivation Path

MEDIUM

Explain the full reasoning path for a linear regression problem with squared loss. Your explanation should begin with empirical risk minimization, identify the role of least squares, state why the objective gradient is compared with zero, interpret A w = b, and finish by distinguishing the invertible and non-invertible cases for A.

Hints
  • Describe what is minimized on the training set.
  • Treat the zero-gradient condition as the bridge from optimization to linear algebra.
  • State which symbols are known from the problem formulation and which vector must be found.
  • Mention eigenvalue decomposition for the non-invertible case.

What do you think happens?

After the squared-loss objective has been formed, what mathematical condition is used to derive the system for the regression weights?

  • Compare the objective's gradient with zero
  • Discard the objective and choose weights arbitrarily
  • Assume that A is invertible before forming a system
Reveal answer

Answer: Compare the objective's gradient with zero

The gradient condition is the step that is rewritten as the linear system A w = b.

The Complete Picture

  1. Empirical risk minimization selects the linear predictor with the smallest training loss; with squared loss, least squares performs this search.
  2. The gradient of the squared-loss objective is compared with zero because that condition is used to derive the weights.
  3. The derived condition is rewritten as the central linear system A w = b, where w contains the regression weights.
  4. If A is invertible, the system has the corresponding direct linear-algebra solution.
  5. If A is not invertible, eigenvalue decomposition provides the needed tools; in the described setting, the system still has a solution because b is in the range of A.

Key Takeaways

  • Least squares is empirical risk minimization for the linear regression class when the loss is squared loss.
  • The zero-gradient condition transforms the optimization problem into a solvable mathematical condition.
  • The central computational form is A w = b, with w representing the regression weights.
  • An invertible A supports the corresponding direct solution, while a non-invertible A calls for eigenvalue-decomposition tools.
  • When A is not invertible in the described setting, b lies in the range of A, so the system still has a solution.