Dimensionality Reduction Techniques
Clustering finds meaningful groups among data objects during exploratory data analysis.
A Dataset Before Its Groups Are Known
When people first examine a dataset, they often want to know whether meaningful groups exist within it. At that stage, the data can be viewed as a set of objects whose groups have not yet been identified. Clustering provides a way to organize those objects by similarity during exploratory data analysis.
The central change is organizational: objects that were initially considered without identified groups are represented as members of groups based on similarity.
The Grouping Decision
Clustering finds meaningful groups among data objects. Similar objects belong in the same group, while dissimilar objects are separated into different groups. A useful way to follow the reasoning is to ask three questions: What are the objects? What kind of similarity matters for this analysis? Which objects should be placed together or separated?
Tracing a Clustering Decision
Suppose an analyst has a collection of data objects and wants to determine whether meaningful groups exist.
Identify the objects: Treat the entries being studied as the objects that may eventually belong to groups.
Choose what similarity matters: Decide which notion of similarity is relevant to the analysis. The appropriate meaning depends on the application.
Place similar objects together: Objects considered similar are represented as members of the same group.
Separate dissimilar objects: Objects considered dissimilar are assigned to different groups.
The previously unorganized objects are represented as groups based on the similarity that matters for the analysis.
Why Exploration Comes First
Clustering is useful in exploratory data analysis because it helps an analyst look for meaningful groups before those groups have been identified. Instead of beginning with known categories, the analyst examines whether the data appears to contain groups according to a relevant idea of similarity. This makes clustering a way to investigate the structure of a dataset.
One Idea Across Many Fields
The same clustering idea can be applied to different kinds of objects. The source identifies genes, customers, and stars as examples, and also notes that clustering is used across the social sciences, biology, computer science, and other disciplines. The objects change from one application to another, but the organizing question remains: which objects should be considered similar for this analysis?
| Application object | Clustering question |
|---|---|
| Genes | Which genes should be considered similar for the analysis? |
| Customers | Which customers should be considered similar for the analysis? |
| Stars | Which stars should be considered similar for the analysis? |
The object type changes across applications, while the clustering task remains the organization of similar objects into groups.
The Similarity Problem
The basic definition of clustering is useful, but it is not completely rigorous by itself. The reason is that similarity does not have one fixed meaning across every application. What counts as similar for genes may differ from what counts as similar for customers or stars. Because the meaning of similarity changes with the application, a fully rigorous general definition of clustering is difficult.
| Question | If the answer changes... |
|---|---|
| What are the objects? | The things being grouped change. |
| What similarity matters? | The basis for grouping changes. |
| Which objects are dissimilar? | The separation between groups changes. |
Treating similarity as having the same meaning in every application.
The meaning of similarity changes with the application.
Fix:
Identify what similarity matters before deciding which objects belong together.Defining clustering only as putting objects into groups.
The definition includes grouping similar objects and separating dissimilar objects.
Fix:
Explain the grouping in relation to similarity and dissimilarity.Assuming that groups are already known before exploration.
Clustering is presented as a technique for finding meaningful groups during exploratory data analysis.
Fix:
Begin with objects whose meaningful groups have not yet been identified.
Check Your Understanding
An analyst is examining a new collection of data objects and wants to use clustering. Describe the three main reasoning steps the analyst should follow, then explain why the result may depend on the application.
Hints
- Start by identifying what counts as an object.
- Ask what kind of similarity matters for the analysis.
- Connect group membership to similarity and separation to dissimilarity.
What do you think happens?
Suppose two analyses use the same set of objects but have different meanings of similarity. Should you expect their groups to be guaranteed to match?
Reveal answer
Answer: No, because similarity depends on the application.
The source explains that the meaning of similarity changes with the application, which makes a fully rigorous general definition of clustering difficult.
The Working Definition
- Clustering finds meaningful groups among data objects during exploratory data analysis.
- Similar objects are placed in the same group, while dissimilar objects are separated into different groups.
- A clustering analysis begins by identifying the objects and deciding what kind of similarity matters.
- Clustering can be applied to genes, customers, stars, and data objects in other disciplines.
- Because similarity changes with the application, the basic definition of clustering is useful but inherently ambiguous.
Key Takeaways
- Clustering organizes data objects into groups according to similarity.
- It is useful for exploratory data analysis because it helps reveal potentially meaningful groups that were not yet identified.
- Genes, customers, stars, and other objects can all be studied with the same general clustering idea.
- The definition becomes ambiguous because the meaning of similarity depends on the application.