Classification with Logistic Regression
Logistic regression is used for classification tasks.
From Input to Probability
Logistic regression is used for classification tasks, but its hypothesis does not produce a class label directly as its first operation. Instead, it transforms an input into a value in the interval [0, 1]. For an input x, this value h(x) can be interpreted as the probability that x has label 1.
The central idea is a two-stage transformation: a linear function first produces a real-valued score, and the logistic sigmoid function then transforms that score into a value in [0, 1].
What do you think happens?
Which stage should receive the original input x first?
Reveal answer
Answer: The linear stage
The linear function maps the input from R^d to a real number. The sigmoid function accepts that real-valued result and maps it into [0, 1], so the order is linear first and sigmoid second.
The Composed Hypothesis
The hypothesis class in logistic regression is formed by composing the logistic sigmoid function over the class of linear functions L_d. In practical terms, choose a linear function, apply it to x, and then pass the resulting real number to the sigmoid function. The complete composition is the logistic regression hypothesis.
Tracing One Input Through the Hypothesis
Describe what happens to an input x as it passes through a logistic regression hypothesis.
Start with x: The input x belongs to the input space R^d.
Apply a linear function: The linear stage maps x from R^d to a real number. This real-valued result is the score produced by the first stage.
Apply the sigmoid function: The sigmoid stage takes the real-valued score as its input and maps it into the interval [0, 1].
Interpret h(x): The final value h(x) is interpreted as the probability that x has label 1.
The composed hypothesis turns an input in R^d into a value in [0, 1] with a class-1 probability interpretation.
Changing Spaces at Each Stage
| Stage | Input | Output | Role |
|---|---|---|---|
| Linear stage | An input x from R^d | A real number | Produces the score |
| Sigmoid stage | A real number | A value in [0, 1] | Produces the final hypothesis value |
The input and output spaces of the two composed stages
Keeping these spaces separate prevents a common confusion. The linear function does not directly produce the final probability interpretation. It produces a real number. The sigmoid function then takes that real number and converts it into the bounded interval [0, 1].
Two Function Classes
| Function class | Receives | Produces | Place in the composition |
|---|---|---|---|
| Linear functions L_d | An input from R^d | A real number | First |
| Logistic sigmoid function | A real number | A value in [0, 1] | Second |
Reading the Final Value
The final hypothesis value h(x) is not described merely as an arbitrary number. Its intended interpretation is the probability that the input x has label 1. The sigmoid stage is what places the result in [0, 1], making this probability interpretation possible.
Interpreting an Abstract Output
Suppose a logistic regression hypothesis produces h(x) for an input x. What should you ask about this value?
Check the range: Because the sigmoid stage maps its input into [0, 1], h(x) belongs to that interval.
Identify the meaning: The value h(x) can be interpreted as the probability that x has label 1.
Keep the stages distinct: The value is produced after the linear score has been passed through the sigmoid function; it is not the raw output of the linear stage.
Read h(x) as a class-1 probability, while remembering that it comes from a linear stage followed by a sigmoid stage.
Mistakes in the Two-Stage Model
Treating the linear output as the final probability.
The linear stage produces a real-valued score, while the sigmoid stage produces the value in [0, 1] used for the class-1 probability interpretation.
Fix:
Trace both stages: first obtain the real-valued score, then apply the sigmoid function to obtain h(x).Reversing the order of the functions.
The source composition places the linear function first. The sigmoid function receives the real number produced by that linear function.
Fix:
Use the order linear stage, then sigmoid stage.Saying that logistic regression immediately outputs a class label.
The hypothesis first produces a value in [0, 1], which is interpreted as the probability that x has label 1.
Fix:
Describe the immediate output as h(x), a value in [0, 1] with a class-1 probability interpretation.Blending the two mappings into one unexplained step.
The linear and sigmoid functions have different input and output spaces and distinct roles.
Fix:
Name the intermediate real-valued score before describing the final value in [0, 1].
Trace It Yourself
For an input x from R^d, write a four-step trace of a logistic regression hypothesis. Your trace should name the input space, the output of the linear stage, the input and output of the sigmoid stage, and the interpretation of h(x).
Hints
- The linear stage comes first.
- The linear stage produces a real number.
- The sigmoid stage maps that real number into [0, 1].
- Interpret the final value in relation to label 1.
Key Takeaways
- Logistic regression is used for classification tasks.
- Its hypothesis class composes the logistic sigmoid function over the class of linear functions L_d.
- The linear stage maps an input from R^d to a real number.
- The sigmoid stage maps that real number into [0, 1].
- The final value h(x) can be interpreted as the probability that x has label 1.