A Clustering Model
The supplied source pack does not define PCA or variance maximization.
A Boundary Before the Details
The title of this article connects clustering with PCA and variance maximization, but the supplied material does not define PCA, variance maximization, principal components, or dimensionality reduction. It does establish clustering: an unsupervised grouping task whose result depends on choices about similarity and distance. This article therefore treats the boundary carefully. It does not explain PCA as if its mechanism were supplied; instead, it develops the clustering model that the material actually supports.
What Enters and Leaves a Clustering Model
A clustering model receives observations together with a way to judge similarity or distance between them. It produces groups, often called clusters. Because the grouping depends on the similarity and distance choices, the model is not determined by the observations alone. Changing those choices can change the resulting clustering.
An Abstract Clustering Input and Output
Suppose a model receives observations and a specified distance rule. What kind of result should you expect from the model?
Input: The model starts with observations and a way to compare their distances or similarities.
Grouping: The model uses those comparisons to decide which observations belong together under the selected clustering procedure.
Output: The result is a set of groups, or cluster assignments. The exact groups depend on the selected similarity or distance choices.
A clustering model maps observations and comparison choices to groups; it does not have one distance-independent answer.
Clustering and PCA Are Not One Task
| Topic | What the supplied material establishes | What is not established |
|---|---|---|
| Clustering | An unsupervised grouping task based on similarity and distance choices | No single result independent of those choices |
| PCA | Only that it is named in the article context | Its procedure, variance maximization, principal components, and dimensionality-reduction behavior |
The Linkage-Based Merge Loop
Linkage-based clustering begins with clusters and repeatedly merges the two clusters that are currently closest. Each merge reduces the number of clusters by one. Repeating this process creates a larger hierarchy of clusters. The procedure therefore has a changing state: the clusters present after one merge become the candidates considered at the next step.
Following Three Linkage Steps
Trace an illustrative linkage-based process that starts with four one-member clusters and repeatedly merges the closest available pair.
Start: Begin with four separate clusters: {A}, {B}, {C}, and {D}.
First merge: The procedure selects the closest pair under its chosen cluster-distance rule and merges that pair. In this illustration, the state becomes {A,B}, {C}, and {D}.
Second merge: The procedure evaluates the current clusters again, rather than continuing to compare only the original individual points. It merges another closest pair, giving {A,B} and {C,D} in this illustration.
Third merge: The remaining two clusters are merged, producing {A,B,C,D}.
The number of clusters falls from four to three, then two, then one. The exact merge order depends on the distance definition and the stopping choice.
Single Linkage and the Closest Pair
Single Linkage defines the distance between two clusters as the minimum distance between any two members of those clusters. To compare cluster X with cluster Y, consider every point in X against every point in Y and use the distance of the closest cross-cluster pair.
Comparing Two Clusters with Single Linkage
Cluster X contains p1 and p2. Cluster Y contains q1 and q2. How does Single Linkage determine the distance between X and Y?
List cross-cluster pairs: Consider the distances between p1 and q1, p1 and q2, p2 and q1, and p2 and q2.
Find the minimum: Identify the smallest distance among those cross-cluster pairs.
Use that pair: The smallest pairwise distance becomes the distance between Cluster X and Cluster Y under Single Linkage.
Single Linkage compares clusters through their closest pair of members, not through every pair equally and not through a representative value supplied by the source.
Single Linkage changes the level at which distance is interpreted. The original distance is between points; Single Linkage turns those point-to-point distances into a cluster-to-cluster distance by selecting the minimum cross-cluster distance.
Two Decisions That Shape the Result
A linkage-based algorithm requires two decisions. First, it must define the distance between clusters because the original input supplies a distance between points. Single Linkage makes one such choice by using the minimum distance between any two members. Second, the algorithm must specify when to stop merging. Different choices for either decision can lead to different clusterings.
| Decision | Question it answers | Why it matters |
|---|---|---|
| Cluster-distance definition | How far apart are two clusters? | It determines which pair is considered closest during a merge. |
| Stopping rule | When should merging stop? | It determines how many clusters remain in the reported result. |
These two decisions are part of specifying a linkage-based clustering algorithm.
Mistakes Beginners Make
Treating clustering and PCA as interchangeable.
The supplied material defines clustering as an unsupervised grouping task but does not define PCA or its mechanism.
Fix:
State only that clustering groups observations according to similarity or distance choices, and explicitly mark PCA details as unsupported by this material.Assuming the closest two original points are always the next clusters to merge.
Linkage-based clustering repeatedly compares the current clusters, and a cluster-distance rule is needed to make that comparison.
Fix:
Recalculate or apply the selected cluster-distance definition to the current clusters at every merge step.Defining Single Linkage as the average distance between two clusters.
Single Linkage uses the minimum distance between any two members of the clusters.
Fix:
List the cross-cluster point pairs and select the smallest distance.Assuming the final number of clusters is determined by merging alone.
The stopping decision is one of the choices that specifies a linkage-based algorithm.
Fix:
Identify the stopping rule before interpreting the resulting cluster assignments.
Check Your Understanding
A linkage-based process currently contains three clusters: {A,B}, {C}, and {D}. Explain what the algorithm must do to choose its next merge, and state what Single Linkage would mean if it compared {A,B} with {C}.
Hints
- The algorithm needs a distance between the current clusters, not only the original individual points.
- For Single Linkage, compare the distances from A to C and from B to C.
- The smaller of those two distances is the distance between {A,B} and {C}.
What do you think happens?
After a linkage-based merge reduces four clusters to three, what happens to the number of clusters after one more merge?
Reveal answer
Answer: It decreases to two.
Each linkage step merges two current clusters into one, reducing the total number of clusters by one.
The Working Model
- Clustering is an unsupervised grouping task whose result depends on similarity and distance choices.
- A clustering model takes observations and comparison choices as input and produces cluster assignments as output.
- Linkage-based clustering repeatedly merges the closest current clusters, reducing the number of clusters at each step.
- Single Linkage defines the distance between two clusters as the minimum distance between any two members.
- The supplied material does not define PCA, variance maximization, principal components, or dimensionality reduction, so those ideas should not be presented as established by this article.
Key Takeaways
- Clustering groups observations using selected similarity or distance choices.
- Linkage-based clustering repeatedly merges the closest current clusters, creating a hierarchy of larger clusters.
- Single Linkage measures two clusters by the closest pair of members across them.
- The distance definition and stopping rule both affect the resulting clustering.
- The supplied material does not define PCA or dimensionality reduction, so those topics must remain outside the article's supported technical claims.