Concepts / Ridge Regression

Ridge Regression

Feature manipulation transforms each original feature to create a resulting feature vector.

  • Programming

From Features to a Learning Input

A learning algorithm does not work directly with an abstract description of a problem. It works with a feature vector: a representation built from the original features. Feature manipulation changes that representation before it is given to the learner. Ridge Regression is connected to this idea because its setup includes feature normalization and a learning rule derived from linear regression.

Feature manipulation is not automatically good or bad. Its value depends on the learning algorithm being used and on the prior assumptions made about the problem.

transformtransformcombinecombineprovide inputOriginal feature 1feature valueTransformed feature 1manipulated valueResulting featurevectorlearner inputLearning algorithmOriginal feature 2feature valueTransformed feature 2manipulated value
How does each original feature get transformed, and how do the transformed features combine into the resulting feature vector used by the learner?

Why Transform Features

Feature manipulation transforms each original feature to create a resulting feature vector. There are two stated purposes for doing this. A transformation may reduce approximation or estimation errors, or it may produce a faster algorithm. These purposes are different: one concerns the errors produced by the learning process, while the other concerns how efficiently the algorithm operates.

Two Possible Reasons for One Transformation

Suppose a learner receives several original features. Why might a practitioner transform those features before learning?

Purpose one: The transformation may be chosen to reduce approximation or estimation errors.

Purpose two: The transformation may instead be chosen to obtain a faster algorithm.

Selection principle: The transformation should be considered together with the learning algorithm and the prior assumptions about the problem. There is no claim that one transformation is universally best.

Feature manipulation is useful only in relation to the problem, its assumptions, and the algorithm that will process the resulting feature vector.

Normalization and the Regression Setup

Normalization is introduced as a type of feature manipulation. Its motivation comes from a linear regression problem using squared loss. In the stated setup, the data is represented by a matrix whose rows are instance vectors and by a vector containing target values.

containsnormalizeprovide featuresprovide targetInstance vectorrow of XFeature valuesoriginal representationRidge RegressioninputNormalized featurestransformed representationTarget valueentry of y
How are an instance vector, its target value, and its normalized feature values related when forming the Ridge Regression input?

The rows of the matrix X are instance vectors, and y is a vector of target values. Normalization changes the feature representation before the learning rule is applied. In this discussion, the normalized features remain part of the regression input, while the target values describe what the learner is trying to use the input to model.

Building the Ridge Rule

Ridge Regression is best understood as a construction. Start with linear regression in the squared-loss setting. Apply the RLM rule with Tikhonov regularization. The resulting learning rule is called Ridge Regression.

useapplyregularize withproduceLinear regressionSquared lossRLM ruleTikhonovregularizationRidge Regression
How do the original linear regression setup, squared loss, RLM rule, and Tikhonov regularization combine step by step to produce the Ridge Regression learning rule?

Classifying a Learning Rule by Construction

A description mentions linear regression, squared loss, the RLM rule, and Tikhonov regularization. What learning rule does this combination identify?

Identify the starting model: The starting problem is linear regression.

Identify the loss: The regression problem uses the squared loss.

Identify the rule: The RLM rule is applied to that problem.

Identify the regularization: Tikhonov regularization is used in obtaining the result.

Name the result: The full combination identifies Ridge Regression.

Ridge Regression is recognized from the complete construction, not from any one ingredient considered alone.

Tikhonov regularization is the technique used in obtaining Ridge Regression. It is not, by itself, the complete name of the resulting learning rule.

Original and Regularized Objectives

The original linear regression setup is described through the model and the squared loss. Ridge Regression is the resulting regularized learning rule after the RLM rule is applied with Tikhonov regularization. The distinction is therefore a distinction between the starting regression problem and the rule produced after regularization is introduced.

usesusesadds regularizationLinear regressionSquared lossRidge RegressionSquared loss withregularization
What changes between the original linear regression objective and the Ridge Regression objective when the regularization term is added?
AspectOriginal linear regressionRidge Regression
Starting pointLinear regressionLinear regression
Loss settingSquared lossSquared loss with regularization
Additional constructionNo stated regularization stepRLM rule with Tikhonov regularization
ResultOriginal regression setupRegularized learning rule called Ridge Regression

Common Classification Mistakes

  • Treating Tikhonov regularization as the complete name of the learning rule.

    Tikhonov regularization is a technique used in obtaining Ridge Regression. The full identification also requires the linear regression setup, squared loss, and RLM rule.

    Fix: Classify the complete combination of model, loss, RLM rule, and regularization technique.

  • Calling every feature transformation universally useful.

    The correct transformation depends on the learning algorithm and the prior assumptions about the problem.

    Fix: Evaluate a transformation in the context of the algorithm that will receive the resulting feature vector.

  • Confusing the original regression setup with the regularized learning rule.

    Ridge Regression is obtained after applying the RLM rule with Tikhonov regularization to the stated linear regression problem.

    Fix: Trace the construction from the starting regression problem to the resulting regularized rule.

  • Ignoring the representation used by the learner.

    The algorithm works with a feature vector built from the original features.

    Fix: Identify the feature manipulation and the resulting feature vector before discussing the learning rule.

Trace the Construction

MEDIUM

A learning problem is represented by a matrix whose rows are instance vectors and by a vector of target values. The features are normalized. The learner uses linear regression with squared loss, applies the RLM rule, and uses Tikhonov regularization. Identify the resulting learning rule and explain the role of each ingredient.

Hints
  • Begin with the model and loss: linear regression with squared loss.
  • Then identify the rule being applied: the RLM rule.
  • Finally identify the regularization technique and use the complete construction to name the result.

Practice Solution

Use the stated ingredients to identify the rule.

Representation: The matrix rows are instance vectors, and the target vector contains target values. Normalization is a feature manipulation applied within this setup.

Starting problem: The model is linear regression with squared loss.

Rule and regularization: The RLM rule is applied with Tikhonov regularization.

Classification: The full construction identifies the resulting learning rule as Ridge Regression.

The answer is Ridge Regression because all of the defining construction steps occur together.

Summary

  1. Feature manipulation transforms original features into a resulting feature vector used by a learning algorithm.
  2. Its purposes can include reducing approximation or estimation errors or obtaining a faster algorithm.
  3. A transformation should be chosen in relation to the learning algorithm and the prior assumptions about the problem.
  4. Normalization is motivated through linear regression with squared loss, where matrix rows are instance vectors and a separate vector contains target values.
  5. Ridge Regression is constructed by applying the RLM rule with Tikhonov regularization to linear regression with squared loss.

Key Takeaways

  • Feature manipulation changes the representation supplied to a learner.
  • Normalization is discussed as feature manipulation within a squared-loss linear regression setting.
  • Ridge Regression is identified by its complete construction: linear regression, squared loss, the RLM rule, and Tikhonov regularization.
  • Tikhonov regularization is a technique in the construction, while Ridge Regression is the resulting learning rule.
  • The key reasoning method is to trace the process from the original regression setup to the regularized rule.