Concepts / Choosing Loss Functions for Classification and Regression

Choosing Loss Functions for Classification and Regression

A Keras functional model can share one representation and branch into several named prediction heads.

  • Programming

One Input, Several Predictions

A model does not need a separate network for every property it predicts. In the Keras functional API, one input can pass through shared layers that build a common internal representation. Near the end, that representation can branch into several named prediction heads. Each head makes a different prediction from the same underlying information.

The source describes a social-media model that uses posts from one person to predict age, income group, and gender at the same time. The shared part processes the input once; the separate heads turn that shared representation into task-specific predictions.

processed bybranches tobranches tobranches toPostsShared representationCommon processing pathAge headScalar regressionIncome headIncome-group classificationGender headBinary classification
How does one input pass through shared layers and then split into separate classification and regression predictions?

Where the Shared Path Splits

The important structural decision is where the model stops being shared. The input and intermediate layers form one common processing path. At the end, separate Dense layers convert the common representation into task-specific predictions. The heads are not interchangeable: each one represents a different kind of target.

Reading the Three Output Heads

Identify the prediction type represented by each head in the source model.

Age: The age head has one unit and treats age as scalar regression.

Income: The income head has one unit per income group and uses softmax, so it represents classification over income groups.

Gender: The gender head has one unit and uses sigmoid, so it represents binary classification.

The same shared representation supports one regression output and two classification outputs.

Pairing Heads with Losses

Because the outputs represent different tasks, training needs a separate loss for each head. A loss should match the kind of prediction and target being evaluated. The source names mean squared error for regression, categorical crossentropy for classification over categories, and binary crossentropy for binary classification.

usesusesusesAgeScalar regressionMSERegression lossIncome groupCategorical classificationCategoricalcrossentropyClassification lossGenderBinary classificationBinary crossentropyBinary classification loss
Which loss function is connected to each output head, and how does that choice depend on whether the output is classification or regression?

From Individual Losses to One Objective

Each output produces its own loss during training. Keras then combines those individual losses into one global loss used for optimization. This lets one training process update a model whose shared representation serves several tasks.

contributescontributescontributesformsAge lossMSELoss combinationKeras combines per-outputlossesGlobal lossUsed for optimizationIncome lossCategorical crossentropyGender lossBinary crossentropy
How do the individual losses from several prediction heads become a single value used to update the model?

A plain combination does not guarantee that every task influences learning equally. If one individual loss is numerically much larger than the others, the shared representation can be optimized mainly for that task. Loss weights change each loss's contribution before the global loss is formed, helping balance outputs whose losses have different numerical scales.

weighted contributionweighted contributionweighted contributionAge lossWeight: adjustedBalanced global lossShared optimizationobjectiveIncome lossWeight: adjustedGender lossWeight: adjusted
How does changing a loss weight alter each output's contribution to the total loss?

When outputs have losses on different numerical scales, inspect the per-output losses rather than looking only at the global loss. If one loss is much larger, consider loss weights so that the shared representation is not dominated by that task.

Matching Targets to Named Outputs

During training, one input array is paired with one target array for each output. Keras accepts those target arrays in either of two forms: a list ordered like the model outputs, or a dictionary keyed by the output names.

ordered pairingnamed pairingpositionpositionpositionoutput nameoutput nameoutput nameModel outputsAge, income, genderTarget listOrder must match outputsAge targetTarget dictionaryKeys name outputsIncome targetGender target
How are target values matched to the correct output when targets are supplied as a list versus a dictionary?
Target formHow matching worksMain concern
ListTargets are matched according to the order of the model outputsThe order must remain correct
DictionaryTargets are matched using output names as keysThe names must identify the intended outputs

Tracing a Multi-Output Bug

A multi-output model is easier to debug when each output is followed through its complete path: prediction head, loss, optional loss weight, and target data. This prevents the model from being treated as one undifferentiated object.

  • Using one loss for every output

    The source model contains scalar regression, categorical classification, and binary classification outputs.

    Fix: Choose a loss suited to each head: MSE for the scalar regression head, categorical crossentropy for the income-group classification head, and binary crossentropy for the binary classification head.

  • Assuming the global loss represents equal task influence

    A plain combination does not guarantee equal influence when the individual losses have different numerical scales.

    Fix: Use loss weights to change each loss's contribution before the global loss is formed.

  • Putting list targets in the wrong order

    List-based target pairing depends on the order of the model outputs.

    Fix: Keep the list order aligned with the outputs, or use a dictionary keyed by output names.

  • Debugging only the final global loss

    The problem may belong to one head's loss, weight, or target pairing.

    Fix: Trace each prediction head to its loss, optional weight, and target data.

Practice the Decision Path

MEDIUM

A model has three named outputs: one predicts a continuous age value, one predicts an income group, and one predicts a binary gender value. Describe the loss type you would connect to each output. Then explain why you might add loss weights and choose either an ordered target list or a target dictionary.

Hints
  • First classify each output as scalar regression, categorical classification, or binary classification.
  • Then connect each task to the corresponding loss named in the source.
  • For target organization, compare the order-dependent list form with the name-based dictionary form.

A Complete Trace

Trace the age output from prediction to optimization in the source model.

Prediction head: The age head is a separate output branch with one unit and represents scalar regression.

Loss: Because age is treated as scalar regression, the head uses mean squared error.

Optional balance: If the age loss has a different numerical scale from the other output losses, a loss weight can change its contribution.

Target pairing: The age target must be supplied in the corresponding position of a target list or under the age output name in a target dictionary.

Global objective: The age loss joins the losses from the other heads to form the global loss used for optimization.

A reliable trace follows one output through its head, matching loss, optional weight, target data, and contribution to the global loss.

Key Takeaways

  • A Keras functional model can process one input through shared layers and branch into several named prediction heads.
  • Each head needs a loss suited to its task, such as MSE for scalar regression, categorical crossentropy for income-group classification, or binary crossentropy for binary classification.
  • Keras combines the individual output losses into one global loss used for optimization.
  • Loss weights help prevent a numerically larger output loss from dominating the shared representation.
  • Targets can be supplied as an output-ordered list or as a dictionary keyed by output names; dictionaries make the pairing explicit.