Concepts / Linear Algebra for Machine Learning

Linear Algebra for Machine Learning

For linear regression with squared loss, least squares solves the empirical risk minimization problem.

  • Programming

From Prediction Errors to Least Squares

Suppose a linear regression model must be selected from a class of possible predictors. The training data determine which predictor is best through empirical risk minimization: choose the predictor whose loss on the training set is smallest. When the loss is squared loss, least squares is the algorithm used to carry out this search for the linear regression class.

evaluate predictorsquare lossaggregateminimizeTraining dataPrediction errorsSquared lossesEmpirical riskLeast squares
How do individual prediction errors become squared losses, and how are those losses aggregated into the empirical risk that least squares minimizes?

The important connection is the objective being minimized. Empirical risk minimization supplies the selection principle: choose the predictor with the smallest training loss. Least squares supplies the procedure for the linear regression class when that loss is squared loss. The result is not a separate objective added after regression; least squares is the way this particular empirical risk minimization problem is carried out.

Why the Gradient Condition Matters

Least squares minimizes the squared-loss objective. To locate an optimal set of regression weights, the derivation calculates the gradient of that objective and compares it with zero. The gradient condition is therefore the mathematical condition used to derive the equations that determine the weights.

evaluate objectiveoptimality conditionCandidate weightsOptimal weightsNonzero gradientZero gradient
How does the optimization condition change when the weights reach a candidate optimum?

Setting the gradient equal to zero is not an optional diagnostic performed after training. In this derivation, it is the step that converts the optimization problem into an equation-solving problem. Once the gradient condition has been formed, it can be rewritten as a linear system.

The Central System A w = b

After the gradient condition is formed, the empirical risk minimization problem is rewritten as the compact linear system A w = b. The vector w is the unknown vector of regression weights. The matrix A and vector b are determined by the problem formulation. Solving the system produces the weights that define the selected linear regression predictor.

determinesdeterminesleft sideunknown vectorright sidesolve for wProblem formulationAsystem matrixA w = blinear systemRegression weightssolution for wwregression weightsbright-hand side
How do the problem formulation, the weight vector, and the target side combine into A w = b?

Following the Derivation

Trace the conceptual route from a squared-loss regression objective to the regression weights.

Start with the objective: Least squares uses the squared-loss objective for the linear regression class.

Calculate the gradient: The gradient of that objective is calculated so the optimization condition can be expressed mathematically.

Compare the gradient with zero: The zero-gradient condition is used to locate an optimal set of weights.

Rewrite the condition: The gradient condition is rewritten as the linear system A w = b.

Solve for w: The solution of the system gives the regression weights.

The regression problem has been transformed from minimizing squared loss into solving the system A w = b.

When A Is Invertible

The next question is whether the matrix A in A w = b is invertible. If A is invertible, the system has the corresponding direct linear-algebra solution for w. In this case, once the gradient condition has been converted into the system, the remaining task is to apply the invertible-matrix case to obtain the regression weights.

use inverse caseuse needed toolsanalyze systemA invertibleA not invertibleDirect solutionsolve for wEigenvaluedecompositionlinear-algebra toolsb in range of Asystem has a solution
What changes in the solution process when A is invertible compared with the non-invertible case?
CaseMain implicationNext reasoning step
A is invertibleThe system has the corresponding direct solution for the weightsSolve the linear system in the invertible-matrix case
A is not invertibleA direct inverse-based solution is not availableUse eigenvalue decomposition as a needed linear-algebra tool

What Non-Invertibility Reveals

If A is not invertible, the direct inverse-based route cannot be used. This situation can occur when the training instances do not span the entire space of R d. The system A w = b still has a solution because b is in the range of A, but the matrix requires a different linear-algebra treatment.

analyze structureseparatesseparatescontributes to solvingrequires analysisA not invertibleEigenvaluedecompositionNonzero eigenvaluesdirections with linearinformationZero eigenvaluesdirections requiringspecial treatmentSolution of A w = bb is in the range of A
What happens to the directions in parameter space when A has zero eigenvalues, and how does eigenvalue decomposition reveal which directions require special treatment?

Eigenvalue decomposition provides the needed linear-algebra tools for the non-invertible case. Its role here is to expose the structure of A so that the system can be analyzed even though A has no direct inverse-based solution. The important distinction is not that the original optimization problem disappears; rather, the same system A w = b must be handled with a decomposition suited to a non-invertible matrix.

Common Reasoning Mistakes

  • Treating empirical risk minimization and least squares as unrelated procedures.

    For linear regression with squared loss, least squares is the algorithm used to carry out the empirical risk minimization search.

    Fix: Connect the objective, loss, and algorithm: squared loss defines the empirical risk, and least squares minimizes it for the linear regression class.

  • Treating the zero-gradient condition as an optional diagnostic.

    The gradient condition is used in the derivation to obtain the system that determines the weights.

    Fix: Begin by calculating the gradient of the least-squares objective and comparing it with zero.

  • Forgetting which symbol is unknown in A w = b.

    A and b are determined by the problem formulation, while w is the vector of regression weights to be found.

    Fix: Read the system as a request to solve for w.

  • Assuming that a non-invertible A means there is no solution.

    In the stated setting, b is in the range of A, so A w = b still has a solution.

    Fix: Move from the direct inverse case to the eigenvalue-decomposition tools for the non-invertible case.

Apply the Derivation

MEDIUM

A least-squares linear regression problem has been formulated, but the derivation has not yet been completed. Describe the sequence of reasoning that connects the objective function to the regression weights. Then explain how your next step changes if A is invertible or not invertible.

Hints
  • Start with the gradient of the squared-loss objective.
  • Use the zero-gradient condition to form A w = b.
  • In the invertible case, use the corresponding direct linear-algebra solution.
  • In the non-invertible case, identify eigenvalue decomposition as the needed tool and remember that b is in the range of A.

Practice Solution Path

Explain the solution path without calculating numerical weights.

Identify the learning problem: The task is empirical risk minimization for a linear regression class using squared loss.

Form the condition: Calculate the objective's gradient and compare it with zero.

Use the central system: Rewrite that condition as A w = b, where w contains the regression weights.

Choose the matrix case: If A is invertible, use the corresponding direct solution. If A is not invertible, use eigenvalue decomposition to analyze and solve the system.

The essential workflow is objective, gradient condition, linear system, and matrix-case-specific solution.

The Complete Mental Model

  1. For linear regression with squared loss, least squares solves the empirical risk minimization problem.
  2. The derivation calculates the gradient of the objective and compares it with zero to locate the optimal weights.
  3. The zero-gradient condition is rewritten as the central linear system A w = b.
  4. A and b are determined by the problem formulation, while w is the regression-weight vector to be found.
  5. If A is invertible, the system has the corresponding direct solution; if A is not invertible, eigenvalue decomposition provides the needed tools, and the system still has a solution because b is in the range of A.

Key Takeaways

  • Least squares is empirical risk minimization for linear regression when the loss is squared loss.
  • Setting the objective gradient equal to zero is the derivation step that identifies the equations for the weights.
  • Those equations take the compact form A w = b.
  • An invertible A leads to the corresponding direct solution, while a non-invertible A requires different linear-algebra tools.
  • Eigenvalue decomposition helps analyze the non-invertible case, where b remains in the range of A and the system still has a solution.