Tikhonov Regularization
Ridge Regression is a learning rule derived from linear regression with the squared loss.
From Linear Regression to Ridge Regression
Ridge Regression is not simply a new name for ordinary linear regression. It is constructed by starting with linear regression, using the squared loss, applying the Regularized Loss Minimization rule, and selecting Tikhonov regularization. The resulting learning rule is called Ridge Regression.
The Derivation Sequence
Read the construction from left to right. Linear regression supplies the starting problem. The squared loss supplies the way empirical error is evaluated. RLM supplies the rule that considers empirical risk together with a regularization function. Tikhonov regularization supplies the particular regularization function. The output of this process is a hypothesis, and in this combination the resulting learning rule is Ridge Regression.
Regularized Loss Minimization
Regularized Loss Minimization, or RLM, is a learning rule that combines empirical risk with a regularization function. It evaluates a hypothesis using both considerations and then outputs a hypothesis.
RLM is a two-part decision rule. The first part is empirical risk, which comes from evaluating the loss on the data. In this derivation, the loss is the squared loss. The second part is a regularization function applied to the model's parameter vector. RLM does not describe a rule that considers only one part; it considers both together before producing a hypothesis.
Tikhonov’s Penalty Function
R(w) = λ ‖w‖₂Tikhonov regularization selects the function R(w) = λ ‖w‖₂. Here, w is the weight vector, ‖w‖₂ is its ℓ2 norm, and λ is a positive scalar. The function therefore has two visible ingredients: the length of the weight vector and a positive scaling factor.
Original and Regularized Objectives
The starting linear regression setup is associated with the squared-loss evaluation of empirical risk. The regularized learning rule changes the decision process by applying RLM and adding a regularization function to the considerations. With Tikhonov regularization, that added function is λ ‖w‖₂. The model remains linear regression, but the learning rule is now the regularized construction identified as Ridge Regression.
| Feature | Original setup | Regularized setup |
|---|---|---|
| Model | Linear regression | Linear regression |
| Loss | Squared loss | Squared loss |
| Decision rule | Empirical-risk evaluation | RLM |
| Regularization function | Not included in the stated starting setup | Tikhonov regularization, R(w) = λ ‖w‖₂ |
| Resulting learning rule | Starting linear regression problem | Ridge Regression |
A Norm Calculation
Evaluating Tikhonov Regularization
Let w = (1, 2, 2) and λ = 3. Calculate the ℓ2 norm of w and then evaluate R(w) = λ ‖w‖₂.
Square the components: The components produce 1², 2², and 2².
Add the squares: The sum is 1² + 2² + 2² = 9.
Take the square root: The ℓ2 norm is ‖w‖₂ = √9 = 3.
Apply λ: Using λ = 3, the regularization value is R(w) = 3 × 3 = 9.
‖w‖₂ = 3 and R(w) = 9.
‖w‖₂ = √(1² + 2² + 2²) = 3
The norm calculation happens before λ is applied. First find the length of the weight vector. Then multiply that length by the positive scalar λ to obtain the value of the Tikhonov regularization function.
Mistakes in Classification
Calling Tikhonov regularization itself Ridge Regression.
Tikhonov regularization is the regularization technique used in obtaining Ridge Regression, not the complete name of the resulting rule.
Fix:
Identify the full construction: linear regression with squared loss, RLM, and Tikhonov regularization produces Ridge Regression.Describing RLM as minimizing only empirical risk.
RLM jointly considers empirical risk and a regularization function.
Fix:
Treat RLM as a two-part decision rule whose output is a hypothesis.Treating λ as the weight vector.
The norm is taken on w, while λ is a positive scalar that scales the norm.
Fix:
Calculate ‖w‖₂ first, then multiply by λ.Stopping after summing the squared components.
The ℓ2 norm requires taking the square root after summing the squares.
Fix:
Use ‖w‖₂ = √(1² + 2² + 2²) = 3.
Check the Construction
A learning setup uses linear regression and the squared loss. It then applies RLM with the regularization function R(w) = λ ‖w‖₂. What learning rule does this construction produce, and what does RLM consider before producing its hypothesis?
Hints
- Trace the ingredients from the starting model to the resulting rule.
- Name both parts considered by RLM.
For w = (1, 2, 2) and λ = 3, calculate ‖w‖₂ and R(w). Write the calculation in two stages: first the norm, then the scaled regularization value.
Hints
- Square, sum, and take the square root.
- Multiply the resulting norm by the positive scalar λ.
What do you think happens?
Which sequence correctly identifies the construction of Ridge Regression?
Reveal answer
Answer: Linear regression, squared loss, RLM with Tikhonov regularization
Ridge Regression is identified by the complete combination of the starting model, the squared loss, the RLM rule, and the Tikhonov regularization choice.
Key Takeaways
- Ridge Regression is constructed from linear regression with the squared loss by applying RLM with Tikhonov regularization.
- RLM combines empirical risk with a regularization function and outputs a hypothesis.
- Tikhonov regularization is R(w) = λ ‖w‖₂, where λ is positive.
- The ℓ2 norm is found by squaring vector components, summing them, and taking the square root.
- For w = (1, 2, 2) and λ = 3, ‖w‖₂ = 3 and R(w) = 9.
Key Takeaways
- Ridge Regression is a construction, not an isolated label: linear regression, squared loss, RLM, and Tikhonov regularization all participate.
- RLM jointly considers empirical risk and a regularization function before producing a hypothesis.
- Tikhonov regularization uses R(w) = λ ‖w‖₂ with positive λ.
- The ℓ2 norm is calculated by squaring, summing, and square-rooting the vector components.
- For the source example w = (1, 2, 2) and λ = 3, the norm is 3 and the regularization value is 9.