Neural Network Outputs
Each output of a multiple-output neural network may have its own loss function.
From Predictions to Learning
A neural network can produce more than one output. When it does, each output can be compared with its corresponding actual output by using its own loss function. This creates several loss values. Before gradient descent can use them, the learning process combines those values into one scalar loss.
One Output at a Time
Begin by following each output independently. An output has three connected parts: the network's prediction, the corresponding target value, and a loss function. The loss function measures the difference between that prediction and the actual output. The result is one loss value for that output.
Tracing Two Outputs
A network produces two outputs. Trace what is measured before the losses are combined.
First output: Compare Output A with Target A using the loss function assigned to that output. This produces Loss A.
Second output: Compare Output B with Target B using the loss function assigned to that output. This produces Loss B.
Before combination: The network now has two separate loss values rather than one shared loss value.
Each output has contributed its own measured difference, ready for the averaging step.
Why One Scalar Matters
Multiple output-specific losses cannot remain as several separate values if they are going to guide gradient descent. Gradient descent requires one scalar loss value. Therefore, the separate losses must be turned into one value that represents the combined result of the output comparisons.
Averaging the Losses
The separate loss values are combined through averaging. For two outputs, first obtain one loss from each output, then calculate their average. That average is the single scalar quantity used for gradient descent.
Turning Two Losses into One
Suppose a multiple-output network produces Loss A equal to 0.2 and Loss B equal to 0.6. How does the network obtain one scalar loss?
Collect the losses: The first output contributes 0.2 and the second output contributes 0.6.
Add the loss values: The combined total is 0.8.
Average them: Divide the total by the two loss values: 0.8 divided by 2 equals 0.4.
The average loss is 0.4, which is one scalar value that can guide gradient descent.
What do you think happens?
A network has two separate loss values, 0.2 and 0.6. What scalar loss results from averaging them?
Reveal answer
Answer: 0.4
Averaging the two losses gives (0.2 + 0.6) divided by 2, which is 0.4.
Mistakes About Multiple Losses
Assuming a multiple-output network must use one shared loss function.
Each output may have its own loss function, so each prediction can be compared with its corresponding actual output separately.
Fix:
Follow each output, its target, and its own loss function independently before combining the resulting loss values.Stopping with several loss values.
Gradient descent requires one scalar loss value rather than several separate values.
Fix:
Average the separate losses to produce the scalar value used for gradient descent.Thinking the average replaces the individual measurements.
The complete path begins with a separate loss for each output, followed by averaging.
Fix:
Preserve the sequence: compare each prediction with its target, obtain each loss, and then average the losses.
Check Your Understanding
A neural network has three outputs. Each output is paired with its corresponding target and produces its own loss value. Describe the steps required to create the single value used by gradient descent.
Hints
- Start with the prediction and target for each output.
- Name what each loss function produces.
- Identify the operation that combines the separate losses.
A complete answer should say that each output is compared with its corresponding target by its own loss function, producing three separate loss values. Those values are then combined through averaging, and the average becomes the one scalar loss used to guide gradient descent.
The Complete Learning Path
- A loss function measures the difference between a prediction and its actual output.
- A multiple-output network may assign a separate loss function to each output.
- The separate loss functions produce multiple loss values.
- Gradient descent requires one scalar loss value rather than several separate values.
- The separate losses are combined through averaging, and the average guides gradient descent.
Key Takeaways
- Each output can be compared with its corresponding target using its own loss function.
- The individual comparisons produce separate loss values.
- Gradient descent requires one scalar loss value.
- Averaging the separate losses creates that scalar value.
- The full path is output and target comparison, individual losses, averaging, and gradient descent.