Exploratory Data Analysis
Clustering organizes objects into groups based on resemblance.
Finding Structure Without Labels
Suppose you receive a collection of objects but no group labels. A natural first question is whether some objects resemble one another more than they resemble the rest. Clustering addresses this question by organizing objects into groups so that similar objects are together and dissimilar objects are separated.
Clustering is an exploratory technique: instead of beginning with a known classification, an analyst uses the grouping to gain an initial understanding of the data.
A Grouping Trace
From Unlabeled Objects to Proposed Groups
Consider a collection containing P, Q, R, S, and T, with no group labels provided. Use resemblance to propose a grouping.
Inspect relationships: Ask whether some objects resemble one another more than they resemble the rest of the collection.
Join similar objects: Place P with Q when they are judged to resemble one another, and place R with S when they are judged to resemble one another.
Separate the remaining object: Keep T in its own group when it does not resemble the other objects closely enough to join them.
A suitable proposed grouping places P with Q, R with S, and T separately.
The objects themselves have not been altered. The change is in how the collection is organized: relationships of resemblance are used to propose groups. This organization can help an analyst recognize structure that was not visible from the unlabeled collection alone.
Meaningful Groups
Clustering is useful when its proposed groups help someone recognize meaningful structure in the data. The purpose is not merely to place every object into a group. In exploratory analysis, the grouping provides an initial way to investigate relationships within a collection that did not come with a known classification.
A clustering result is valuable when the organization of similar objects makes the data easier to understand or investigate.
Clustering Across Fields
| Field | Objects being grouped | Similarity described in the source |
|---|---|---|
| Biology | Genes | Similarities in expression across experiments |
| Retail | Customers | Customer profiles |
| Astronomy | Stars | Spatial proximity |
These applications use different kinds of objects and observations, but they share the same purpose: applying clustering to identify meaningful groups. Computational biologists can group genes according to similarities in their expression across experiments. Retailers can group customers using customer profiles to support targeted marketing. Astronomers can group stars according to spatial proximity.
The Similarity Problem
The grouping idea therefore does not determine every clustering result by itself. An analyst still has to address what counts as resemblance for the data being explored and where group boundaries should be placed. The source definition identifies this unresolved issue without giving one universal answer.
Common Interpretation Mistakes
Treating clustering as a known classification
Clustering is commonly used when group labels are not provided and the analyst is exploring the data.
Fix:
Treat the groups as proposed organization intended to support an initial understanding of the data.Assuming clustering changes the objects
The important change is how the collection is organized, not a change to the objects themselves.
Fix:
Describe clustering as using resemblance relationships to propose groups.Assuming that similarity has one automatically precise meaning
The basic definition leaves the precise meaning of similarity and dissimilarity unresolved.
Fix:
Ask what resemblance means for the particular data set and recognize that the intuitive definition is not fully precise.
Practice the Definition
Explain clustering in two parts: first describe how similarity determines group membership, then explain why the resulting groups are useful during exploratory data analysis.
Hints
- Mention that similar objects are organized together and dissimilar objects are separated.
- Mention that clustering helps an analyst gain an initial understanding of data without beginning with known classification labels.
- Include the limitation that the definition does not specify exactly how similarity must be defined for every data set.
Classifying Three Applications
Match each situation with the objects being grouped and the resemblance used: genes across experiments, customers for targeted marketing, and stars in space.
Biology: Genes are grouped according to similarities in their expression across experiments.
Retail: Customers are grouped using customer profiles to support targeted marketing.
Astronomy: Stars are grouped according to spatial proximity.
Each situation uses clustering to identify meaningful groups, but the objects and observations used to judge resemblance differ by field.
Key Takeaways
- Clustering organizes objects into groups based on resemblance.
- It is commonly used to explore data and discover meaningful structure without beginning with known classification labels.
- The objects themselves are not altered; their organization is changed by proposing groups based on relationships of resemblance.
- Examples include grouping genes, customers, and stars using different kinds of observations.
- The intuitive definition is not fully precise because it does not specify exactly how similarity, dissimilarity, and group boundaries must be defined for every data set.
Key Takeaways
- Clustering groups objects that resemble one another and separates objects judged to be dissimilar.
- It helps analysts explore unlabeled data and recognize meaningful structure.
- Genes, customers, and stars are examples of objects that can be grouped in different fields.
- The definition is intuitive but incomplete because the precise meaning of similarity and group boundaries remains unresolved.