Sigmoid Function
Logistic regression is used for classification tasks.
From Input to Probability
Logistic regression is used for classification tasks, but its hypothesis does not begin by producing a class label directly. Instead, it transforms an input into a value in the interval [0, 1]. For an input x, this value h(x) can be interpreted as the probability that x has label 1.
The Two Function Stages
The hypothesis class in logistic regression is formed by composing the logistic sigmoid function over the class of linear functions L_d. This means that logistic regression uses two function classes in a fixed order: a linear function first and the sigmoid function second.
| Stage | Input space | Output space | Role |
|---|---|---|---|
| Linear function | R^d | The real numbers | Maps the input to a real-valued score |
| Sigmoid function | The real numbers | [0, 1] | Maps the score to a bounded value |
The input and output spaces of the two stages
Tracing a Composed Hypothesis
Following One Input Through Both Functions
Trace an input x through the two stages of a logistic regression hypothesis.
Start with the input: The input x belongs to the input space R^d.
Apply the linear function: The linear stage maps x to a real number. This intermediate result is a real-valued score, not yet a probability-like output.
Apply the sigmoid function: The sigmoid stage receives that real number and maps it into the interval [0, 1].
Interpret h(x): The final value h(x) is interpreted as the probability that x has label 1.
The composed hypothesis first produces a real-valued score and then converts that score into a value in [0, 1] with a probability interpretation.
Reading the Final Value
The final hypothesis value is not described as a direct class label. It is a value in [0, 1], and for an input x, h(x) can be interpreted as the probability that x has label 1. The probability interpretation belongs to the completed composition, after the real-valued result from the linear stage has passed through the sigmoid stage.
Generated example: If a completed logistic regression hypothesis produces h(x) = 0.8 for an input x, the value is interpreted as an 0.8 probability that x has label 1. The number 0.8 is being used as an illustration of a value in [0, 1]; the important point is the interpretation of the final hypothesis value, not the particular number.
Common Interpretation Errors
Treating the linear-stage output as the final probability.
The linear stage produces an intermediate real-valued score. The probability interpretation belongs to the value after the sigmoid stage.
Fix:
Trace both stages: first the linear mapping, then the sigmoid mapping into [0, 1].Putting the sigmoid function before the linear function.
The stated composition uses the linear function first. The sigmoid function accepts the real-valued result of that first stage.
Fix:
Use the order input from R^d, linear function to a real number, then sigmoid function to [0, 1].Assuming the hypothesis directly produces a class label as its first operation.
A logistic regression hypothesis first produces a value in [0, 1]. That value can be interpreted as the probability that the input has label 1.
Fix:
Distinguish the final probability-like hypothesis value from a direct class label.Blending the two mappings into one unexplained step.
This hides the separate roles of the linear function and the sigmoid function.
Fix:
Name the two mappings separately: R^d to the real numbers, followed by the real numbers to [0, 1].
Check Your Understanding
An input x is passed through a logistic regression hypothesis. Describe the input and output of each stage, state which stage comes first, and explain how the final value h(x) should be interpreted.
Hints
- Start with the input space R^d.
- The first stage produces a real number.
- The second stage produces a value in [0, 1].
- The final value is interpreted with respect to label 1.
What do you think happens?
Before reading the explanation, predict what kind of value the linear stage produces and what kind of value the sigmoid stage produces.
Reveal answer
Answer: The linear stage produces a real number; the sigmoid stage produces a value in [0, 1].
The linear stage maps the input from R^d to a real number. The sigmoid stage then maps that real number into [0, 1], where the final value can be interpreted as the probability that the input has label 1.
Key Takeaways
- Logistic regression is used for classification tasks.
- Its hypothesis class composes the logistic sigmoid function over the class of linear functions L_d.
- The linear stage maps an input from R^d to a real number.
- The sigmoid stage maps that real number into [0, 1].
- The final value h(x) can be interpreted as the probability that x has label 1.
Key Takeaways
- Logistic regression combines a linear function with a logistic sigmoid function.
- The linear function comes first and maps R^d to a real number.
- The sigmoid function comes second and maps that real number into [0, 1].
- The completed hypothesis value h(x) can be interpreted as the probability that x has label 1.