Domain Points and Labels
Training data is the learner’s input and consists of a finite sequence of labeled examples.
From Inputs to Training Data
Before a learner can use the information in a learning problem, it needs input. Training data is that input. It is not merely a collection of described objects: it is a finite sequence of examples in which every input is paired with a label.
A training-data pair has exactly two identified roles: an input from X and a label from Y.
Reading One Labeled Pair
The formal representation of training data is S = ((x₁, y₁), ..., (xₘ, yₘ)). Each parenthesized pair is one labeled training example. In every pair, xᵢ is the domain point, or input, and yᵢ is its label. The notation X × Y expresses that each pair contains one element from X together with one element from Y.
Identifying the Parts of an Example
Consider one pair written as (x₁, y₁). Identify its input and its label.
First part: x₁ is the domain point or input. It comes from X.
Second part: y₁ is the label associated with that input. It comes from Y.
Whole pair: Together, (x₁, y₁) is one labeled training example.
The first coordinate is the input from X, and the second coordinate is the label from Y.
A Sequence of Examples
Training data contains more than one possible example: it is a finite sequence of labeled examples. The complete sequence is written as S, and its individual entries are the pairs (x₁, y₁) through (xₘ, yₘ). The pairs are the training examples; together, they make up the complete training data.
Suppose described papayas are used as inputs. Color and softness describe each papaya, while tastiness is the result associated with that described example. A described papaya together with its tastiness label forms one labeled pair. Several such pairs form the training data.
Names for the Collection
The words training data and training set refer to the complete collection S. A training example refers to one individual labeled pair inside that collection. This gives a useful level distinction: one pair is an example, while all the pairs together are the training set or training data.
Labels Versus Inputs Alone
A collection containing only inputs has domain points but does not show the labels associated with those points. Labeled training data has a pair for each example: an input from X together with a label from Y. The structural difference is therefore the presence or absence of the second part of each example.
| Collection | Each entry contains | Status in this concept |
|---|---|---|
| Labeled training data | An input from X and a label from Y | Training data |
| Collection of inputs without labels | An input from X only | Missing the label part of a training-data pair |
Common Identification Mistakes
Calling one pair the entire training set.
A training example is one individual labeled pair, while the training set or training data is the complete collection.
Fix:
Use training example for one pair and training set or training data for the complete sequence S.Treating an input by itself as a labeled example.
A labeled example requires both an input from X and a label from Y.
Fix:
Check that every example is represented as a pair (xᵢ, yᵢ).Assuming that a larger collection of inputs is automatically training data.
The collection becomes training data when the domain information is paired with labels.
Fix:
Look for the label associated with each described input.Confusing the domains X and Y.
The input comes from X, while the label comes from Y.
Fix:
In each pair, identify the first coordinate as the input from X and the second as the label from Y.
Check Your Understanding
A collection is written as S = ((x₁, y₁), (x₂, y₂), (x₃, y₃)). Identify one training example, identify the complete training data, and state what would be missing if only x₁, x₂, and x₃ were recorded.
Hints
- A training example is one parenthesized pair.
- The complete training data is the whole sequence named S.
- Compare an entry containing xᵢ and yᵢ with an entry containing xᵢ alone.
Checking the Sequence
Use S = ((x₁, y₁), (x₂, y₂), (x₃, y₃)) to distinguish an example from the complete collection.
Choose one pair: (x₂, y₂) is one training example because it is one input-label pair.
Identify the collection: S is the complete training data, also called the training set.
Remove the labels: If only x₁, x₂, and x₃ remain, the entries no longer display the labels from Y, so the labeled-pair structure is absent.
One pair is a training example; the complete sequence S is the training data or training set; inputs alone omit the label part.
Key Takeaways
- Training data is the learner's input and is a finite sequence of labeled examples.
- Each labeled example is a pair containing an input from X and a label from Y.
- The individual pairs are training examples; the complete collection S is the training data or training set.
- A collection of inputs without labels does not have the labeled-pair structure described by S.
Key Takeaways
- Training data is a finite sequence of labeled examples.
- Each example has two parts: an input from X and a label from Y.
- One pair is a training example, while the full sequence S is the training set or training data.
- Inputs without their associated labels do not form the labeled training-data structure.