Linear Functions
Logistic regression is used for classification tasks.
From Input to Probability
Logistic regression is used for classification tasks, but its hypothesis does not produce a class label directly as its first operation. Instead, it transforms an input into a value in the interval [0, 1]. For an input x, this final value h(x) can be interpreted as the probability that x has label 1.
The central mechanism is a composition: a linear function acts first, and the logistic sigmoid function acts second.
Tracing the Two Stages
The first stage belongs to the class of linear functions L_d. It receives an input from R^d and maps that input to a real number. This real-valued result is the intermediate score produced by the linear stage.
The second stage is the logistic sigmoid function. It receives the real number produced by the linear function and maps that number into the interval [0, 1]. The final result is the hypothesis value h(x), which can be interpreted as the probability that the original input x has label 1.
What do you think happens?
Which stage must come first if the sigmoid function accepts a real number as its input?
Reveal answer
Answer: The linear function
The linear stage maps the input from R^d to a real number. That real-valued result is then suitable as the input to the sigmoid stage, which maps it into [0, 1].
A Composed Hypothesis
A composed hypothesis is formed when the output of one function becomes the input to another function. In logistic regression, the hypothesis class is formed by composing the logistic sigmoid function over the class of linear functions L_d.
Following One Input Through the Composition
Trace a single input x through the two stages of a logistic regression hypothesis.
Start with x: The input x belongs to the input space R^d.
Apply the linear stage: A linear function from L_d receives x and returns one real number.
Pass the intermediate result onward: The real number returned by the linear stage becomes the input to the logistic sigmoid function.
Apply the sigmoid stage: The sigmoid function maps that real number into the interval [0, 1].
Interpret h(x): The resulting value h(x) is interpreted as the probability that x has label 1.
The complete hypothesis is a two-stage mapping from x in R^d to a value h(x) in [0, 1].
Reading the Final Value
The final value is not the original input and it is not the intermediate real number from the linear stage. It is the output of the sigmoid stage. Because the sigmoid maps into [0, 1], this output can be interpreted as a probability: h(x) is the probability that x has label 1.
Interpreting h(x)
A composed hypothesis receives an input x and produces h(x) in [0, 1]. What should the learner understand from that output?
Identify the stage: The value h(x) is produced after the sigmoid stage, not directly after the linear stage.
Identify the range: The sigmoid output lies in [0, 1].
Apply the interpretation: The value is interpreted as the probability that the input x has label 1.
The hypothesis value h(x) is a probability interpretation for label 1, produced by applying the sigmoid to the linear stage's real-valued output.
Mistakes in the Two-Stage Model
Treating the linear function as the complete logistic regression hypothesis.
The linear stage only maps an input from R^d to a real number. The probability interpretation comes after the sigmoid stage.
Fix:
Trace the input through both stages: linear function first, sigmoid function second.Reversing the order of the functions.
The sigmoid stage receives a real number produced by the linear stage.
Fix:
Remember that the linear output is the sigmoid input.Saying that the hypothesis immediately returns a class label.
The hypothesis first produces a value in [0, 1], which can be interpreted as the probability that x has label 1.
Fix:
Describe h(x) first as a probability-valued output.Confusing the input and output spaces of the stages.
The linear stage maps from R^d to a real number, while the sigmoid stage maps from a real number to [0, 1].
Fix:
Name the spaces separately for each stage.
Check Your Understanding
Explain the complete path of an input x through a logistic regression hypothesis. Your explanation must name the input space of x, the output of the linear stage, the input and output of the sigmoid stage, and the interpretation of h(x).
Hints
- Start with x in R^d.
- The linear stage returns a real number.
- That real number becomes the sigmoid input.
- The final value lies in [0, 1] and is interpreted as the probability of label 1.
A learner says, "Logistic regression uses a sigmoid function to map an input from R^d directly to a probability." Rewrite the statement so that it correctly describes the two-stage composition.
Hints
- Identify the stage that accepts the input from R^d.
- Identify the intermediate value passed between the stages.
- State what the final value represents.
Key Takeaways
- Logistic regression is used for classification tasks.
- Its hypothesis class composes the logistic sigmoid function over the class of linear functions L_d.
- The linear stage maps an input from R^d to a real number.
- The sigmoid stage maps that real number into [0, 1].
- The final hypothesis value h(x) can be interpreted as the probability that x has label 1.
Key Takeaways
- Logistic regression combines a linear function with a logistic sigmoid function.
- The linear function acts first, mapping an input from R^d to a real number.
- The sigmoid function acts second, mapping that real number into [0, 1].
- The final value h(x) is interpreted as the probability that the input has label 1.