The Learner in Statistical Learning
Training data is the learner’s input and consists of a finite sequence of labeled examples.
The Learner’s Input
Before a learner can use information in a learning problem, it needs input. In statistical learning, that input is training data: a finite sequence of labeled examples. Each example connects a domain point with the label associated with it.
The word labeled is essential. A collection of inputs becomes training data when each input is paired with a label.
Building the Sequence
Training data is written as S = ((x₁, y₁), ..., (xₘ, yₘ)). This notation shows that the learner receives a finite sequence with m labeled examples. The parentheses group each input and its label into one example, while the outer sequence collects all examples into S.
Reading a Training Sequence
Interpret S = ((x₁, y₁), (x₂, y₂), (x₃, y₃)) as training data.
Find the examples: The three parenthesized pairs are the individual training examples: (x₁, y₁), (x₂, y₂), and (x₃, y₃).
Separate each pair: In every pair, xᵢ is the input or domain point, and yᵢ is the label associated with that input.
Identify the complete data: The entire sequence is S, so S is the complete training data and can also be called the training set.
The sequence contains three labeled training examples, and the complete collection is the training set S.
Inside One Labeled Example
A training-data pair is an element of X × Y. Its first component, xᵢ, is a domain point from X. Its second component, yᵢ, is the corresponding label from Y.
The pairing is what connects information about an example to the result associated with that example. The input alone describes the domain point, but the label supplies the associated outcome needed for labeled learning data.
Consider papayas described by properties such as color and softness. Those properties provide information about each described papaya, while tastiness is the result associated with it. The collection becomes training data when each description is paired with its tastiness label.
Labels Versus Inputs
| Collection | What each item contains | Is it labeled training data? |
|---|---|---|
| Labeled training data | An input from X paired with a label from Y | Yes |
| Collection of inputs | Inputs without their associated labels | No |
From Examples to Training Set
A training example is one labeled pair, such as (xᵢ, yᵢ). Training data is the complete finite sequence of those pairs. Training set is another name for the complete training data S, not the name of one individual pair.
- One pair is one training example.
- A finite sequence of pairs is the training data.
- The complete training data S is also called the training set.
Common Terminology Mistakes
Calling an input by itself a labeled training example.
A training example is a pair containing both an input and its label.
Fix:
Write the complete example as (xᵢ, yᵢ).Calling a collection of inputs training data without checking for labels.
Training data consists of labeled examples, not inputs alone.
Fix:
Verify that every input is paired with its corresponding label.Using training set to mean one example.
The training set is another name for the complete training data S.
Fix:
Use training example for one pair and training set for the full sequence.
When reading a description of a learning problem, identify three levels separately: the input xᵢ, the labeled pair (xᵢ, yᵢ), and the complete sequence S. This prevents the terms example, training data, and training set from being mixed together.
Check Your Understanding
A learning problem contains a finite collection of domain points. For each point, a corresponding label is also recorded. Explain whether the collection is training data, identify the form of one training example, and state what the complete collection can be called.
Hints
- Ask whether each domain point is paired with a label.
- Represent one example as (xᵢ, yᵢ).
- Use the term training set for the complete training data.
What do you think happens?
Suppose a collection contains x₁, x₂, and x₃, but no y values. Is it already labeled training data?
Reveal answer
Answer: No, because the inputs are not paired with labels.
Training data requires a finite sequence of labeled examples, and each example contains an input paired with a label.
Key Takeaways
- Training data is the learner’s input and is a finite sequence of labeled examples.
- Each training example is a pair (xᵢ, yᵢ), with xᵢ from X and yᵢ from Y.
- The complete sequence is written as S = ((x₁, y₁), ..., (xₘ, yₘ)).
- Training set is another name for the complete training data S.
- Inputs without associated labels form a collection of inputs, not labeled training data.
Key Takeaways
- Training data is the learner’s input: a finite sequence of labeled examples.
- A labeled example contains an input from X and a label from Y.
- The notation S = ((x₁, y₁), ..., (xₘ, yₘ)) represents the complete training data.
- Individual pairs are training examples, while the complete collection is the training set.
- A collection of inputs is not labeled training data until labels are paired with those inputs.