Concepts / Sigmoid Function

Sigmoid Function

Logistic regression is used for classification tasks.

  • Programming

From Input to Probability

Logistic regression is used for classification tasks, but its hypothesis does not begin by producing a class label directly. Instead, it transforms an input into a value in the interval [0, 1]. For an input x, this value h(x) can be interpreted as the probability that x has label 1.

entersmaps toentersmaps toInput xR^dLinear functionL_dReal-valued scorea real numberSigmoid functioninput: real numberh(x)[0, 1]
How does an input move through the linear function and then the sigmoid function to produce the final hypothesis value?

The Two Function Stages

The hypothesis class in logistic regression is formed by composing the logistic sigmoid function over the class of linear functions L_d. This means that logistic regression uses two function classes in a fixed order: a linear function first and the sigmoid function second.

StageInput spaceOutput spaceRole
Linear functionR^dThe real numbersMaps the input to a real-valued score
Sigmoid functionThe real numbers[0, 1]Maps the score to a bounded value

The input and output spaces of the two stages

linear functionsigmoid functionR^dinput spaceReal numberlinear output[0, 1]sigmoid output
What type of value enters and leaves the linear stage, and how does the sigmoid stage transform that value into a number between 0 and 1?

Tracing a Composed Hypothesis

Following One Input Through Both Functions

Trace an input x through the two stages of a logistic regression hypothesis.

Start with the input: The input x belongs to the input space R^d.

Apply the linear function: The linear stage maps x to a real number. This intermediate result is a real-valued score, not yet a probability-like output.

Apply the sigmoid function: The sigmoid stage receives that real number and maps it into the interval [0, 1].

Interpret h(x): The final value h(x) is interpreted as the probability that x has label 1.

The composed hypothesis first produces a real-valued score and then converts that score into a value in [0, 1] with a probability interpretation.

producesbecomes input toproducesLinear functionoutput: real numberReal-valued scoreintermediate resultSigmoid functionoutput: [0, 1]h(x)probability of label 1
What is the difference between the linear function's score and the sigmoid function's probability-like output, and how are they connected?

Reading the Final Value

The final hypothesis value is not described as a direct class label. It is a value in [0, 1], and for an input x, h(x) can be interpreted as the probability that x has label 1. The probability interpretation belongs to the completed composition, after the real-valued result from the linear stage has passed through the sigmoid stage.

Generated example: If a completed logistic regression hypothesis produces h(x) = 0.8 for an input x, the value is interpreted as an 0.8 probability that x has label 1. The number 0.8 is being used as an illustration of a value in [0, 1]; the important point is the interpretation of the final hypothesis value, not the particular number.

interpreted ash(x)value in [0, 1]Probability of label1interpretation for input x
How should the final sigmoid output be interpreted as the probability that the input belongs to label 1?

Common Interpretation Errors

  • Treating the linear-stage output as the final probability.

    The linear stage produces an intermediate real-valued score. The probability interpretation belongs to the value after the sigmoid stage.

    Fix: Trace both stages: first the linear mapping, then the sigmoid mapping into [0, 1].

  • Putting the sigmoid function before the linear function.

    The stated composition uses the linear function first. The sigmoid function accepts the real-valued result of that first stage.

    Fix: Use the order input from R^d, linear function to a real number, then sigmoid function to [0, 1].

  • Assuming the hypothesis directly produces a class label as its first operation.

    A logistic regression hypothesis first produces a value in [0, 1]. That value can be interpreted as the probability that the input has label 1.

    Fix: Distinguish the final probability-like hypothesis value from a direct class label.

  • Blending the two mappings into one unexplained step.

    This hides the separate roles of the linear function and the sigmoid function.

    Fix: Name the two mappings separately: R^d to the real numbers, followed by the real numbers to [0, 1].

Check Your Understanding

EASY

An input x is passed through a logistic regression hypothesis. Describe the input and output of each stage, state which stage comes first, and explain how the final value h(x) should be interpreted.

Hints
  • Start with the input space R^d.
  • The first stage produces a real number.
  • The second stage produces a value in [0, 1].
  • The final value is interpreted with respect to label 1.

What do you think happens?

Before reading the explanation, predict what kind of value the linear stage produces and what kind of value the sigmoid stage produces.

  • The linear stage produces a real number; the sigmoid stage produces a value in [0, 1].
  • The linear stage produces a value in [0, 1]; the sigmoid stage produces an input in R^d.
  • Both stages directly produce class labels.
Reveal answer

Answer: The linear stage produces a real number; the sigmoid stage produces a value in [0, 1].

The linear stage maps the input from R^d to a real number. The sigmoid stage then maps that real number into [0, 1], where the final value can be interpreted as the probability that the input has label 1.

Key Takeaways

  1. Logistic regression is used for classification tasks.
  2. Its hypothesis class composes the logistic sigmoid function over the class of linear functions L_d.
  3. The linear stage maps an input from R^d to a real number.
  4. The sigmoid stage maps that real number into [0, 1].
  5. The final value h(x) can be interpreted as the probability that x has label 1.

Key Takeaways

  • Logistic regression combines a linear function with a logistic sigmoid function.
  • The linear function comes first and maps R^d to a real number.
  • The sigmoid function comes second and maps that real number into [0, 1].
  • The completed hypothesis value h(x) can be interpreted as the probability that x has label 1.