Nearest-Neighbor Selection
k-NN regression predicts a real-valued target by combining the targets of the k nearest neighbors.
From Neighbors to a Prediction
Suppose a new input x does not yet have a known target. The k-nearest-neighbors method looks at the training examples closest to x. In a regression problem, the target space is the set of real numbers, written as Y = R. The method uses the target values attached to the k nearest neighbors to produce one real-valued prediction.
The prediction does not come from choosing one neighbor's target. It combines the targets of all k selected neighbors.
The central question is therefore not only which point is closest. It is also how the selected target values are combined after the neighbors have been identified.
Selecting the Nearest Neighbors
For a new input x, the training examples are ordered by their distance from x. The position of a neighbor in this ordering is represented by π_i(x), where i identifies its position. The first k positions identify the k selected neighbors. Their input locations determine selection, while their attached target values are used later to calculate the prediction.
Choosing Three Neighbors
A new input is compared with training points ordered from closest to farthest. If k = 3, which points contribute to the prediction?
Order the points: Use the distance ordering around the new input. The first point is π₁(x), the second is π₂(x), and the third is π₃(x).
Apply k: Because k = 3, select the first three positions in the ordering.
Exclude the rest: Points appearing after the first three do not contribute to this particular k-nearest-neighbors prediction.
The selected neighbors are π₁(x), π₂(x), and π₃(x).
Averaging the Selected Targets
The standard regression rule computes an ordinary average. Add the target values of the k selected neighbors and divide by k. If h_S(x) denotes the prediction for x, then the prediction uses the target associated with each selected neighbor π_i(x). Every selected target has equal participation in this ordinary average.
Computing an Ordinary k-NN Average
The three selected neighbors have target values 10, 20, and 30. What prediction does the ordinary k-NN regression rule produce?
Identify k: There are three selected targets, so k = 3.
Add the targets: The total is 10 + 20 + 30 = 60.
Divide by k: Divide the total by 3: 60 divided by 3 equals 20.
The regression prediction is 20.
With the ordinary average rule, changing the order of the selected targets does not change the result, because each selected target contributes equally.
Changing k Changes the Result
The value of k controls how many neighbor targets participate in the prediction. A smaller k uses fewer targets, while a larger k uses more targets. Because the selected targets are the values being averaged, changing k can change both the selected set and the resulting prediction.
Comparing Two Values of k
The ordered neighbor targets are 10, 20, 30, and 50. Compare the ordinary-average predictions for k = 3 and k = 4.
Use k = 3: Average the first three targets: 10 + 20 + 30 = 60, and 60 divided by 3 equals 20.
Use k = 4: Include the fourth target as well: 10 + 20 + 30 + 50 = 110, and 110 divided by 4 equals 27.5.
Compare: The fourth target participates when k increases from 3 to 4, so the average changes.
The prediction is 20 for k = 3 and 27.5 for k = 4.
The Generalized Function φ
The ordinary average is one rule for combining the selected neighbor information. A generalized k-NN rule uses a function φ to map the selected neighbor input-target pairs to an output target. In this view, neighbor selection happens first, and φ specifies what to do with the selected pairs afterward.
When φ is the ordinary average rule, it adds the selected target values and divides by k. Other choices of φ can define different combination rules. For example, a distance-based weighted average gives closer neighbors more influence instead of giving every selected neighbor equal influence.
| Rule | What is combined | Influence of selected neighbors |
|---|---|---|
| Ordinary average | The k selected target values | Equal participation |
| Distance-based weighted average | The k selected target values using a distance-based rule | Closer neighbors receive more influence |
Common Selection Mistakes
Using only the single closest target when k is greater than 1.
The standard regression rule uses all k selected targets in the average.
Fix:
Identify the first k neighbors in the distance ordering, then combine all k target values.Selecting neighbors by their target values rather than by their distance from the new input.
Nearest-neighbor selection is based on which training inputs are closest to x. The target values are used after selection.
Fix:
Order the training inputs by distance from x before reading the attached targets.Changing k but keeping the old prediction.
Changing k changes how many target values participate, so the average may change.
Fix:
Rebuild the selected set and recompute the combination rule whenever k changes.Assuming every k-NN rule must be an ordinary average.
The generalized function φ can specify other ways to map selected neighbor pairs to an output target.
Fix:
Identify the rule represented by φ before deciding how the selected targets contribute.
Practice the Prediction
A new input has four nearest neighbors with target values ordered by distance as 6, 12, 18, and 30. Using the ordinary average rule, calculate the prediction for k = 2 and for k = 4. Then explain why the two predictions differ.
Hints
- For k = 2, use only the first two target values in the distance ordering.
- For k = 4, use all four target values.
- For each value of k, add the selected targets and divide by k.
Practice Solution
Use the ordered target values 6, 12, 18, and 30 to calculate the ordinary-average predictions for k = 2 and k = 4.
Calculate k = 2: Use 6 and 12. Their sum is 18, and 18 divided by 2 equals 9.
Calculate k = 4: Use 6, 12, 18, and 30. Their sum is 66, and 66 divided by 4 equals 16.5.
Explain the difference: The k = 4 calculation includes two additional target values, so it produces a different average.
The prediction is 9 for k = 2 and 16.5 for k = 4.
Key Takeaways
- k-NN regression predicts a real-valued target from the targets attached to the k nearest neighbors.
- The standard rule adds the k selected targets and divides by k, giving every selected target equal participation.
- The value of k determines how many neighbor targets contribute, so changing k can change the prediction.
- The function φ describes a generalized rule for mapping selected neighbor input-target pairs to an output target.
- A distance-based weighted average is different from an ordinary average because closer neighbors receive more influence.
Key Takeaways
- Select neighbors according to their distance from the new input.
- For ordinary k-NN regression, average the target values of the k selected neighbors.
- Increasing or decreasing k changes the set of targets that participates in the prediction.
- The generalized function φ represents the rule used to combine selected neighbor input-target pairs.
- Equal participation describes the ordinary average, while distance-based weighting gives closer neighbors more influence.