Concepts / Modeling Relationships Between Variables

Modeling Relationships Between Variables

Linear regression is used to model relationships between explanatory variables and a real-valued outcome.

  • Programming

From Information to an Estimate

Suppose you want to estimate a baby's weight. You might use information such as her age and weight at birth. The information used to make the estimate consists of explanatory variables. The quantity being estimated is the outcome. Linear regression is a statistical tool for modeling this kind of relationship.

helps estimatehelps estimateAgeexplanatory variableBaby's weightreal-valued outcomeWeight at birthexplanatory variable
Which variables are used as inputs, and which single real-valued quantity is the model trying to predict?

The direction of the problem matters: explanatory variables provide the input, while the real-valued outcome is what the learned predictor aims to approximate.

The Learning Problem

When linear regression is viewed as a learning problem, its input domain X is a subset of R^d for some d. An input is therefore represented as a point in a d-dimensional real-valued space. Its label set Y is the set of real numbers. The outcome is a real-valued label associated with an input.

containspaired withcontainsXsubset of R^dxone inputYreal numbersyone real-valued outcome
How does an input from the domain map to a real-valued label in the learning problem?

Linear regression is used to model relationships between explanatory variables and a real-valued outcome.

Choosing a Function

The goal is to learn a linear function h that maps an input from R^d to a real number. The function h is a predictor: it takes the explanatory-variable input and produces an estimated outcome. The learning goal is not to choose any possible kind of function. It is to find a function that belongs to the allowed family and best approximates the relationship between the inputs and the outcome.

A hypothesis class is the collection of candidate functions that a learning problem is allowed to consider. In linear regression, the hypothesis class is restricted to linear functions.

includesincludesmay be selectedmay be selectedLinear functionshypothesis classCandidate function AlinearSuitable memberlearned predictorCandidate function Blinear
What family of possible functions can the learner choose from before selecting a suitable predictor?

Following the Predictor

Estimating a Baby's Weight

Identify the inputs, outcome, domain, label set, and learned predictor in a linear-regression learning problem about estimating a baby's weight.

Identify the explanatory variables: The baby's age and weight at birth are information used to make the estimate, so they serve as explanatory variables.

Identify the outcome: The baby's weight is the quantity being estimated, so it is the real-valued outcome.

Describe the domain: The combined input belongs to the input domain X, which is a subset of R^d for some d.

Describe the label set: The outcome belongs to Y, the set of real numbers.

Describe the learned predictor: The learning task is to learn a linear function h that maps the input to a real number and best approximates the relationship between the explanatory variables and the baby's weight.

The explanatory variables form the input, the baby's weight is the real-valued outcome, and the learned linear predictor maps the input to an estimated real number.

describeapproximated bymaps input toInput-outcome pairsexplanatory variables andoutcomeRelationshipto be approximatedEstimated outcomereal numberhlearned linear predictor
How does the learned predictor approximate the underlying relationship between explanatory variables and the observed outcome?

The learned predictor is the result of the modeling process. It is not the explanatory variable and it is not the observed outcome. It is a learned linear function that uses the input to produce a real-number estimate, with the aim of approximating the relationship represented by the data.

Mistakes in Problem Direction

  • Treating the outcome as an explanatory variable

    The explanatory variables provide the input, while the real-valued outcome is what the predictor aims to approximate.

    Fix: First identify what information is used to make the estimate. Then identify the single quantity being estimated.

  • Describing linear regression as only a description of data

    Linear regression is also a way to learn a predictor from the relationship being modeled.

    Fix: Include the learned function h and its role in mapping an input to a real number.

  • Assuming the hypothesis class contains every possible function

    The hypothesis class in linear regression is restricted to linear functions.

    Fix: State that the learner selects or learns a suitable member of the linear-function class.

  • Confusing the domain with the label set

    The input domain X is a subset of R^d for some d, while the label set Y is the set of real numbers.

    Fix: Keep the input space X and the real-valued label set Y distinct when describing the learning problem.

Before discussing a linear model, state the direction of the problem in words: these explanatory variables are the input, and this real-valued quantity is the outcome. This simple step prevents the model's inputs, labels, and learned predictor from being mixed together.

Practice and Summary

MEDIUM

A learning problem uses several explanatory variables to estimate one real-valued quantity. Explain which part is the input, which part is the outcome, what X and Y represent, and why the learned predictor belongs to the linear-function hypothesis class.

Hints
  • Start by identifying the information used to make the estimate.
  • Then identify the quantity being estimated.
  • Use X for the input domain and Y for the real-valued label set.
  • Finish by explaining that the learned predictor is a suitable linear function.
  1. Linear regression models relationships between explanatory variables and a real-valued outcome. Its input domain X is a subset of R^d for some d, and its label set Y is the set of real numbers. The hypothesis class contains linear functions. The learning goal is to learn a suitable linear function h that maps an input to a real number and best approximates the relationship between the inputs and the outcome.

Key Takeaways

  • Explanatory variables provide the input to a linear-regression problem.
  • The outcome is a real-valued quantity that the predictor aims to approximate.
  • The input domain X is a subset of R^d for some d, while the label set Y is the set of real numbers.
  • The hypothesis class in linear regression consists of linear functions.
  • A learned predictor h maps an input to a real number and approximates the relationship between the explanatory variables and the outcome.