Concepts / Classification with Logistic Regression

Classification with Logistic Regression

Logistic regression is used for classification tasks.

  • Programming

From Input to Probability

Logistic regression is used for classification tasks, but its hypothesis does not produce a class label directly as its first operation. Instead, it transforms an input into a value in the interval [0, 1]. For an input x, this value h(x) can be interpreted as the probability that x has label 1.

The central idea is a two-stage transformation: a linear function first produces a real-valued score, and the logistic sigmoid function then transforms that score into a value in [0, 1].

What do you think happens?

Which stage should receive the original input x first?

  • The sigmoid stage
  • The linear stage
  • Either stage, because the order does not matter
Reveal answer

Answer: The linear stage

The linear function maps the input from R^d to a real number. The sigmoid function accepts that real-valued result and maps it into [0, 1], so the order is linear first and sigmoid second.

The Composed Hypothesis

inputproducesbecomes inputproducesInput xR^dLinear functionL_dReal-valued scoreSigmoid functionlogistic sigmoidh(x)[0, 1]
How does the linear function feed its output into the sigmoid function to form the final hypothesis?

The hypothesis class in logistic regression is formed by composing the logistic sigmoid function over the class of linear functions L_d. In practical terms, choose a linear function, apply it to x, and then pass the resulting real number to the sigmoid function. The complete composition is the logistic regression hypothesis.

Tracing One Input Through the Hypothesis

Describe what happens to an input x as it passes through a logistic regression hypothesis.

Start with x: The input x belongs to the input space R^d.

Apply a linear function: The linear stage maps x from R^d to a real number. This real-valued result is the score produced by the first stage.

Apply the sigmoid function: The sigmoid stage takes the real-valued score as its input and maps it into the interval [0, 1].

Interpret h(x): The final value h(x) is interpreted as the probability that x has label 1.

The composed hypothesis turns an input in R^d into a value in [0, 1] with a class-1 probability interpretation.

Changing Spaces at Each Stage

inputmaps toinputmaps toR^dinput spaceLinear stageL_dReal numberscore spaceSigmoid stagelogistic sigmoid[0, 1]probability range
What does each stage take as input and produce as output, and how do those spaces change from feature vector to real-valued score to probability?
StageInputOutputRole
Linear stageAn input x from R^dA real numberProduces the score
Sigmoid stageA real numberA value in [0, 1]Produces the final hypothesis value

The input and output spaces of the two composed stages

Keeping these spaces separate prevents a common confusion. The linear function does not directly produce the final probability interpretation. It produces a real number. The sigmoid function then takes that real number and converts it into the bounded interval [0, 1].

Two Function Classes

Linear functionR^d to real numberSigmoid functionreal number to [0, 1]
What is the difference between the linear function that produces a score and the sigmoid function that transforms that score?
Function classReceivesProducesPlace in the composition
Linear functions L_dAn input from R^dA real numberFirst
Logistic sigmoid functionA real numberA value in [0, 1]Second

Reading the Final Value

sigmoid transformationReal-valued scorebefore sigmoidh(x)probability of label 1 in[0, 1]
How does the final hypothesis value represent the probability that an input belongs to label 1?

The final hypothesis value h(x) is not described merely as an arbitrary number. Its intended interpretation is the probability that the input x has label 1. The sigmoid stage is what places the result in [0, 1], making this probability interpretation possible.

Interpreting an Abstract Output

Suppose a logistic regression hypothesis produces h(x) for an input x. What should you ask about this value?

Check the range: Because the sigmoid stage maps its input into [0, 1], h(x) belongs to that interval.

Identify the meaning: The value h(x) can be interpreted as the probability that x has label 1.

Keep the stages distinct: The value is produced after the linear score has been passed through the sigmoid function; it is not the raw output of the linear stage.

Read h(x) as a class-1 probability, while remembering that it comes from a linear stage followed by a sigmoid stage.

Mistakes in the Two-Stage Model

  • Treating the linear output as the final probability.

    The linear stage produces a real-valued score, while the sigmoid stage produces the value in [0, 1] used for the class-1 probability interpretation.

    Fix: Trace both stages: first obtain the real-valued score, then apply the sigmoid function to obtain h(x).

  • Reversing the order of the functions.

    The source composition places the linear function first. The sigmoid function receives the real number produced by that linear function.

    Fix: Use the order linear stage, then sigmoid stage.

  • Saying that logistic regression immediately outputs a class label.

    The hypothesis first produces a value in [0, 1], which is interpreted as the probability that x has label 1.

    Fix: Describe the immediate output as h(x), a value in [0, 1] with a class-1 probability interpretation.

  • Blending the two mappings into one unexplained step.

    The linear and sigmoid functions have different input and output spaces and distinct roles.

    Fix: Name the intermediate real-valued score before describing the final value in [0, 1].

Trace It Yourself

EASY

For an input x from R^d, write a four-step trace of a logistic regression hypothesis. Your trace should name the input space, the output of the linear stage, the input and output of the sigmoid stage, and the interpretation of h(x).

Hints
  • The linear stage comes first.
  • The linear stage produces a real number.
  • The sigmoid stage maps that real number into [0, 1].
  • Interpret the final value in relation to label 1.

Key Takeaways

  • Logistic regression is used for classification tasks.
  • Its hypothesis class composes the logistic sigmoid function over the class of linear functions L_d.
  • The linear stage maps an input from R^d to a real number.
  • The sigmoid stage maps that real number into [0, 1].
  • The final value h(x) can be interpreted as the probability that x has label 1.