Concepts / Real-Valued Prediction

Real-Valued Prediction

Linear regression is used to model relationships between explanatory variables and a real-valued outcome.

  • Programming

From Information to an Estimate

Many prediction problems begin with information about something and ask you to estimate a quantity. For example, you might want to estimate a baby's weight using information such as her age and weight at birth. The information used to make the estimate is made up of explanatory variables. The quantity being estimated is the outcome. When that outcome is a number on the real-number scale, linear regression provides a statistical tool for modeling the relationship.

inputinputpredictsAgeExplanatory variableLearned predictorLinear functionBaby's weightReal-valued outcomeBirth weightExplanatory variable
How do explanatory variables flow into a learned predictor that produces a real-valued outcome?

Input and Outcome Roles

The direction of the modeling problem matters. Explanatory variables provide the input to the prediction process. The real-valued outcome is the quantity that the predictor aims to estimate. In the baby-weight example, age and weight at birth are explanatory variables, while the baby's weight is the outcome.

Linear regression is a statistical tool used to model relationships between explanatory variables and a real-valued outcome.

Sorting the Roles in a Prediction Problem

Suppose the goal is to estimate a baby's weight using her age and weight at birth. Identify the explanatory variables and the outcome.

Identify the information available for prediction: The baby's age and weight at birth are the information supplied to the prediction process. They are therefore explanatory variables.

Identify the quantity being estimated: The baby's weight is the quantity the model is asked to estimate, so it is the outcome.

Identify the type of prediction: Because the outcome is a real-valued quantity, this is a real-valued prediction problem modeled using linear regression.

Age and weight at birth are explanatory variables; the baby's weight is the real-valued outcome.

Domain and Label Set

Linear regression can be described as a learning problem with an input domain and a label set. The input domain X is a subset of R^d for some d. This expresses that each input consists of d real-valued components, with the particular allowed inputs collected in X. The label set Y is the set of real numbers, because the outcome being predicted is real-valued.

containscontainspaired withpaired withXSubset of R^dInputExplanatory variablesInput-outcome pairOne input paired with oneoutcomeYReal numbersOutcomeReal-valued label
What does the domain contain, what does the label set contain, and how are inputs paired with outcomes?

The domain describes the kinds of inputs the learning problem accepts. The label set describes the possible outcomes. For linear regression, the inputs lie in X, a subset of R^d, and the labels lie in Y, the real numbers.

The Linear Hypothesis Class

A hypothesis class is the collection of candidate functions that a learning problem is allowed to consider. In linear regression, the hypothesis class is restricted to linear functions. The learning task is to select or learn a suitable member of this class for approximating the relationship represented by the data.

possible choicepossible choicepossible choiceLinear function ACandidateLearned predictorSelected linear functionLinear function BCandidateLinear function CCandidate
What different linear functions could be considered, and how does one learned function differ from the other candidates?

The phrase linear function hypothesis class emphasizes a restriction: the learner does not consider every possible function. It considers linear functions and learns one that best approximates the relationship between the explanatory variables and the outcome. The learned predictor is therefore one member of the allowed class, chosen for the modeling task.

From Relationship to Predictor

What do you think happens?

If the goal is to estimate a real-valued outcome from explanatory variables, what should the learned function produce?

  • A real number
  • A member of a restricted set of linear-function candidates
  • A real number produced by a learned linear function
  • All of the above
Reveal answer

Answer: All of the above

The learned predictor is a linear function selected from the hypothesis class. It maps an input from R^d to a real number and aims to approximate the relationship between the explanatory variables and the outcome.

modeled byrelationship to approximateproducesObserved inputExplanatory variablesLearned predictorLinear functionObserved outcomeReal-valued outcomePredicted outcomeReal number
How does a learned linear predictor approximate the relationship between observed explanatory-variable values and real-valued outcomes?

The purpose of learning is not merely to describe the variables. A learned function h maps an input from R^d to a real number. Its role is to approximate the relationship between the explanatory variables and the outcome. This is why linear regression can be understood both as a statistical modeling tool and as a way to learn a predictor.

Common Modeling Mistakes

  • Treating every variable as an outcome

    Those quantities provide the information used to make the estimate; they are explanatory variables.

    Fix: Identify the target quantity first. The baby's weight is the outcome in this example.

  • Forgetting that the outcome is real-valued

    For linear regression, the label set Y is the set of real numbers.

    Fix: State that the predicted outcome belongs to the real numbers.

  • Confusing a hypothesis class with one learned function

    A hypothesis class is the collection of candidate functions. The learned predictor is one suitable member of that collection.

    Fix: Use hypothesis class for the allowed collection and learned predictor for the selected linear function.

  • Describing linear regression only as data description

    The learning goal is to learn a linear function that best approximates the relationship between inputs and outcome.

    Fix: Explain both roles: modeling the relationship and learning a predictor.

Check Your Understanding

MEDIUM

A prediction problem uses several explanatory variables to estimate one quantity that belongs to the real numbers. Explain how you would describe its domain, label set, hypothesis class, and learned predictor.

Hints
  • Describe the domain using X and R^d.
  • Describe the label set using Y and the real numbers.
  • Remember that the hypothesis class contains candidate linear functions.
  • Describe the learned predictor as the function that approximates the relationship between inputs and outcome.
  1. To analyze a real-valued prediction problem, first separate the explanatory variables from the outcome. Then describe the inputs as belonging to X, a subset of R^d, and the outcomes as belonging to Y, the real numbers. Linear regression restricts the hypothesis class to linear functions. Learning selects a suitable linear function h that maps an input to a real number and best approximates the relationship between the explanatory variables and the outcome.

Key Takeaways

  • Linear regression models relationships between explanatory variables and a real-valued outcome.
  • The domain X is a subset of R^d, while the label set Y is the set of real numbers.
  • A hypothesis class is a collection of candidate functions; linear regression uses a class of linear functions.
  • The learned predictor is a selected linear function that maps inputs to real numbers.
  • The predictor aims to approximate the relationship between the explanatory variables and the outcome.