Training a Machine Learning Model
A model receives samples and produces predictions; targets provide the external truth used for comparison.
From Input to Learning Signal
Training a machine learning model begins with samples. A sample is an item of data entering the model. The model processes that sample and produces a prediction, which is the result produced by the model. A target provides the external truth: the value the model should ideally have produced. Comparing the prediction with the target produces a loss value, which measures the distance between them.
The prediction and target have different origins. The prediction comes from the model; the target comes from an external source of data. The loss value describes the distance between them.
Vocabulary of a Single Sample
A sample, or input, is one item of data given to the model. A prediction, or output, is the result produced by the model for that sample. A target is the value the model should ideally have produced according to an external source of data. A ground-truth annotation is the externally supplied annotation used as that target. A loss value measures the distance between the prediction and the target, so it provides a way to describe prediction error.
Tracing One Labeled Sample
Suppose a model receives one sample and produces a prediction. An external data source supplies the value that the model should have produced.
Sample: The sample is the data entering the model.
Prediction: The model produces an output for that sample. This output is the prediction.
Target: The externally supplied value is the target, or ground-truth annotation, for the sample.
Loss value: The prediction and target are compared. Their distance is represented by a loss value.
The model output and the external target are not interchangeable: one is produced by the model, while the other supplies the reference for judging that output.
Categories and Assigned Labels
In classification, a class is a possible category. A label identifies the category or categories assigned to a particular sample. The important distinction is whether each sample must receive one of two exclusive categories, one of more than two categories, or multiple labels at the same time.
| Task type | Categories per sample | What the label does |
|---|---|---|
| Binary classification | One of two exclusive categories | Identifies one selected category |
| Multiclass classification | One of more than two categories | Identifies one selected category |
| Multilabel classification | Multiple categories may apply at the same time | Identifies multiple assigned categories |
Consider three generated task descriptions. Deciding between two exclusive outcomes illustrates binary classification. Selecting one category from more than two possible categories illustrates multiclass classification. Assigning several categories to the same sample illustrates multilabel classification. The defining issue is not merely the number of names in the vocabulary; it is how many categories a sample can receive.
Continuous Targets
Regression uses continuous target values rather than choosing from a set of classification labels. The distinction between scalar and vector regression depends on how many continuous values the target contains.
| Regression type | Target contents | Model output |
|---|---|---|
| Scalar regression | One continuous value | One continuous value |
| Vector regression | Multiple continuous values | Multiple continuous values |
Counting Continuous Outputs
A regression task has a target containing either one continuous value or several continuous values. Identify the task type in each case.
One value: A target containing one continuous value is used in scalar regression.
Several values: A target containing multiple continuous values is used in vector regression.
Compare the output: The model output follows the same distinction: one continuous value for scalar regression and multiple continuous values for vector regression.
Count the continuous values in the target. One means scalar regression; multiple means vector regression.
Mini-Batches in Training
A mini-batch, also called a batch, is a small set of samples processed simultaneously by a model. The source describes a typical mini-batch as containing between 8 and 128 samples. The number of samples is often a power of 2 to facilitate memory allocation on a GPU.
A mini-batch changes the unit being processed: instead of presenting one sample at a time, the model processes a small set simultaneously. Each sample still has its own prediction and target relationship, and the resulting loss values describe the prediction distances for the samples in that set. The mini-batch therefore provides a group of training cases for one training step.
Mistakes That Blur the Vocabulary
Calling the model's prediction the target.
The prediction is produced by the model; the target is supplied externally.
Fix:
Use prediction for the model output and target for the external reference value.Treating a loss value as another prediction.
The loss value measures the distance between those two values rather than being the model's output for the sample.
Fix:
Describe the loss as a measure of prediction error.Using multiclass and multilabel as synonyms.
The possibility of multiple labels for one sample is the defining difference.
Fix:
Ask whether one sample can receive several categories simultaneously.Defining regression by the number of classes.
Regression is distinguished from classification by the kind of target, not by a classification category count.
Fix:
For regression, count the continuous values in the target: one for scalar and multiple for vector regression.Assuming a mini-batch is one sample with many labels.
Batch size counts samples, while labels describe categories assigned to an individual sample.
Fix:
Keep the data-grouping idea of a batch separate from the category-assignment idea of labels.
When describing a training task, name the objects in order: sample, prediction, target, and loss value. Then state whether the target is a classification label or a continuous regression value. This vocabulary makes the model's input, output, reference value, and prediction error distinct.
Check Your Understanding
A task assigns exactly one category from four possible categories to each sample. Is it binary, multiclass, or multilabel classification? Then consider a regression target containing three continuous values. Is the regression scalar or vector? Finally, explain what changes when eight samples are processed simultaneously instead of one sample.
Hints
- For classification, ask how many categories are available and whether more than one can be assigned to a sample.
- For regression, count the continuous values in the target.
- For the final question, use the definition of a mini-batch.
Practice Answer
Classify the three situations in the practice prompt.
Four possible categories, one assigned: This is multiclass classification because one category is selected from more than two possible categories.
Three continuous target values: This is vector regression because the target contains multiple continuous values.
Eight samples together: The eight samples form a mini-batch because a mini-batch is a small set of samples processed simultaneously.
The answers are multiclass classification, vector regression, and a mini-batch of eight samples.
The Training Vocabulary
- A sample enters the model, and a prediction leaves it.
- A target, or ground-truth annotation, is the externally supplied value the model should ideally have produced.
- A loss value measures the distance between a prediction and its target.
- Binary, multiclass, and multilabel classification differ in how many categories a sample can receive and whether those categories are exclusive.
- Scalar regression has one continuous target value, vector regression has multiple continuous target values, and a mini-batch is a small set of samples processed simultaneously.
Key Takeaways
- A model receives samples and produces predictions; targets provide the external truth used for comparison.
- A loss value measures the distance between a prediction and its target.
- Classification may be binary, multiclass, or multilabel depending on the number and exclusivity of labels assigned to each sample.
- Regression may be scalar or vector depending on whether the target contains one continuous value or multiple continuous values.
- A mini-batch is a small set of samples processed simultaneously by a model.