Linear Algebra for Machine Learning
For linear regression with squared loss, least squares solves the empirical risk minimization problem.
From Prediction Errors to Least Squares
Suppose a linear regression model must be selected from a class of possible predictors. The training data determine which predictor is best through empirical risk minimization: choose the predictor whose loss on the training set is smallest. When the loss is squared loss, least squares is the algorithm used to carry out this search for the linear regression class.
The important connection is the objective being minimized. Empirical risk minimization supplies the selection principle: choose the predictor with the smallest training loss. Least squares supplies the procedure for the linear regression class when that loss is squared loss. The result is not a separate objective added after regression; least squares is the way this particular empirical risk minimization problem is carried out.
Why the Gradient Condition Matters
Least squares minimizes the squared-loss objective. To locate an optimal set of regression weights, the derivation calculates the gradient of that objective and compares it with zero. The gradient condition is therefore the mathematical condition used to derive the equations that determine the weights.
Setting the gradient equal to zero is not an optional diagnostic performed after training. In this derivation, it is the step that converts the optimization problem into an equation-solving problem. Once the gradient condition has been formed, it can be rewritten as a linear system.
The Central System A w = b
After the gradient condition is formed, the empirical risk minimization problem is rewritten as the compact linear system A w = b. The vector w is the unknown vector of regression weights. The matrix A and vector b are determined by the problem formulation. Solving the system produces the weights that define the selected linear regression predictor.
Following the Derivation
Trace the conceptual route from a squared-loss regression objective to the regression weights.
Start with the objective: Least squares uses the squared-loss objective for the linear regression class.
Calculate the gradient: The gradient of that objective is calculated so the optimization condition can be expressed mathematically.
Compare the gradient with zero: The zero-gradient condition is used to locate an optimal set of weights.
Rewrite the condition: The gradient condition is rewritten as the linear system A w = b.
Solve for w: The solution of the system gives the regression weights.
The regression problem has been transformed from minimizing squared loss into solving the system A w = b.
When A Is Invertible
The next question is whether the matrix A in A w = b is invertible. If A is invertible, the system has the corresponding direct linear-algebra solution for w. In this case, once the gradient condition has been converted into the system, the remaining task is to apply the invertible-matrix case to obtain the regression weights.
| Case | Main implication | Next reasoning step |
|---|---|---|
| A is invertible | The system has the corresponding direct solution for the weights | Solve the linear system in the invertible-matrix case |
| A is not invertible | A direct inverse-based solution is not available | Use eigenvalue decomposition as a needed linear-algebra tool |
What Non-Invertibility Reveals
If A is not invertible, the direct inverse-based route cannot be used. This situation can occur when the training instances do not span the entire space of R d. The system A w = b still has a solution because b is in the range of A, but the matrix requires a different linear-algebra treatment.
Eigenvalue decomposition provides the needed linear-algebra tools for the non-invertible case. Its role here is to expose the structure of A so that the system can be analyzed even though A has no direct inverse-based solution. The important distinction is not that the original optimization problem disappears; rather, the same system A w = b must be handled with a decomposition suited to a non-invertible matrix.
Common Reasoning Mistakes
Treating empirical risk minimization and least squares as unrelated procedures.
For linear regression with squared loss, least squares is the algorithm used to carry out the empirical risk minimization search.
Fix:
Connect the objective, loss, and algorithm: squared loss defines the empirical risk, and least squares minimizes it for the linear regression class.Treating the zero-gradient condition as an optional diagnostic.
The gradient condition is used in the derivation to obtain the system that determines the weights.
Fix:
Begin by calculating the gradient of the least-squares objective and comparing it with zero.Forgetting which symbol is unknown in A w = b.
A and b are determined by the problem formulation, while w is the vector of regression weights to be found.
Fix:
Read the system as a request to solve for w.Assuming that a non-invertible A means there is no solution.
In the stated setting, b is in the range of A, so A w = b still has a solution.
Fix:
Move from the direct inverse case to the eigenvalue-decomposition tools for the non-invertible case.
Apply the Derivation
A least-squares linear regression problem has been formulated, but the derivation has not yet been completed. Describe the sequence of reasoning that connects the objective function to the regression weights. Then explain how your next step changes if A is invertible or not invertible.
Hints
- Start with the gradient of the squared-loss objective.
- Use the zero-gradient condition to form A w = b.
- In the invertible case, use the corresponding direct linear-algebra solution.
- In the non-invertible case, identify eigenvalue decomposition as the needed tool and remember that b is in the range of A.
Practice Solution Path
Explain the solution path without calculating numerical weights.
Identify the learning problem: The task is empirical risk minimization for a linear regression class using squared loss.
Form the condition: Calculate the objective's gradient and compare it with zero.
Use the central system: Rewrite that condition as A w = b, where w contains the regression weights.
Choose the matrix case: If A is invertible, use the corresponding direct solution. If A is not invertible, use eigenvalue decomposition to analyze and solve the system.
The essential workflow is objective, gradient condition, linear system, and matrix-case-specific solution.
The Complete Mental Model
- For linear regression with squared loss, least squares solves the empirical risk minimization problem.
- The derivation calculates the gradient of the objective and compares it with zero to locate the optimal weights.
- The zero-gradient condition is rewritten as the central linear system A w = b.
- A and b are determined by the problem formulation, while w is the regression-weight vector to be found.
- If A is invertible, the system has the corresponding direct solution; if A is not invertible, eigenvalue decomposition provides the needed tools, and the system still has a solution because b is in the range of A.
Key Takeaways
- Least squares is empirical risk minimization for linear regression when the loss is squared loss.
- Setting the objective gradient equal to zero is the derivation step that identifies the equations for the weights.
- Those equations take the compact form A w = b.
- An invertible A leads to the corresponding direct solution, while a non-invertible A requires different linear-algebra tools.
- Eigenvalue decomposition helps analyze the non-invertible case, where b remains in the range of A and the system still has a solution.