Playground / Feature Scaling: Min-Max and Standardization

See how feature scales change the nearest neighbour

Feature Scaling: Min-Max and Standardization

Interactive lab

Try it: Feature Scaling: Min-Max and Standardization

How centering, min-max scaling and standardization change each feature's location and range, and why features on very different scales let one feature dominate Euclidean distances — so the nearest neighbour can change once the features are scaled. The scaler's statistics come from the training data and are then applied unchanged to new data.

How it works

  1. Fit the scaler per feature on the training points: mean (centering), min and max (min-max), or mean and standard deviation (standardization).
  2. Transform every training point with that feature's formula: x − mean, (x − min)/(max − min) (optionally mapped to [−1, 1]), or (x − mean)/std.
  3. Apply the same fitted transform to the test point — it may land outside [0, 1]; the scaler is not refitted on new data.
  4. Measure the Euclidean distance from the test point to every training point in the transformed space and rank them.
  5. Compare with the raw ranking: without scaling, the feature with the larger numbers dominates; centering alone changes nothing.

Default run (11 steps): 6 training points and one test point (100 sq m, 6 rooms). Area spans 40–180, rooms 1.5–7. Method: standardization, fitted on the training points only. … Nearest neighbour after standardization: p_2 (class B); ranking p_2 < p_4 < p_3 < p_6 < p_1 < p_5. With raw values it would be p_1 (class A) — the large-scale area feature dominated.

Simplified: Two features (area 20–200 sq m, rooms 1–8) and at most 8 training points on a grid; the effect is shown on a 1-nearest-neighbour lookup only. Standard deviation is the population value (ddof = 0) and a zero range or zero spread is replaced by 1, as scikit-learn does.

Educational simulation

Loading the simulation…