Clustering Objectives
Mutual information I(X;Y) measures dependence between two random variables.
The Compression–Retention Tension
A clustering method replaces many individual data points with a smaller description: a cluster identity. This replacement is useful only when it removes unnecessary detail while preserving information that matters for the task. For document clustering, the information bottleneck approach describes this tension using mutual information.
The objective has two competing purposes. The first is compression: the cluster representation should not preserve every document-specific detail. The second is retention: the cluster representation should preserve information associated with words, because words are treated as the relevant information about the document.
Symbols in the Objective
Mutual information I(X;Y) measures dependence between two random variables. The variables being compared determine what relationship the quantity describes.
| Symbol | Role in document clustering | Relationship used |
|---|---|---|
| X | Identity of a document | Appears in I(X;C) |
| C | Clustering variable whose values are available cluster labels | Appears in both I(X;C) and I(C;Y) |
| Y | Identity of a word | Appears in I(C;Y) |
| β | Weight on the word-related mutual information term | Controls the contribution of I(C;Y) in the objective |
The symbols identify the document, its cluster, the word, and the tradeoff weight.
The notation changes according to the pair being compared. I(X;C) concerns document identity and cluster identity. I(C;Y) concerns cluster identity and word identity. These pairings serve different purposes, so they should not be read as interchangeable measurements.
Reading the Objective
I(X;C) − βI(C;Y)Interpreting the two terms
Explain what each part of the information bottleneck objective asks a document clustering method to do.
Read I(X;C): This term measures the connection between a document's identity and its cluster identity. Making it small expresses compression: the cluster representation should not preserve all original document-specific detail.
Read I(C;Y): This term measures the connection between the clustering variable and word identity. Retaining high mutual information here preserves information associated with words.
Read the minus sign: Because the word-related term is subtracted, increasing retained word information supports a smaller value of the combined objective, while the first term is being made small directly.
Read β: The weight β determines how strongly the word-related mutual information contributes to the combined expression. It therefore participates in the tradeoff between compression and retention.
The objective favors a cluster representation that compresses document information while retaining information about words.
From Assignments to Clusters
The expression becomes a clustering objective because the method searches over probabilistic point-to-cluster assignments. An assignment describes how documents are associated with the available cluster labels. Each candidate assignment determines the relationships measured by I(X;C) and I(C;Y), and the combined objective evaluates that candidate.
Comparing two candidate assignments
Suppose two different probabilistic assignments of documents to the available cluster labels are being considered. What does the objective-based comparison involve?
Keep the variables fixed: Both candidates use X for document identity, C for the clustering variable, and Y for word identity.
Evaluate the first candidate: The candidate assignment determines one set of document–cluster and cluster–word relationships, which are represented by I(X;C) and I(C;Y).
Evaluate the second candidate: The second assignment determines another set of relationships and therefore another value of the combined objective.
Compare the expressions: The method compares the evaluated objectives and searches over such probabilistic assignments for the minimizing choice.
The cluster labels are produced through an objective-based search over assignments, not merely by naming the two relationships.
Dependence and Information
Mutual information is not a label for one particular kind of variable. It is a measure applied to a pair of random variables. In this objective, changing the pair changes the meaning of the measurement: I(X;C) concerns document and cluster identities, whereas I(C;Y) concerns cluster and word identities.
Common Interpretation Errors
Treating X, C, and Y as interchangeable symbols.
The first pair concerns document identity and cluster identity. The second concerns cluster identity and word identity.
Fix:
Translate each pair before interpreting the term: X with C means document–cluster dependence; C with Y means cluster–word dependence.Describing both terms as compression.
The source assigns compression to making I(X;C) small, while high I(C;Y) represents retention of relevant word-related information.
Fix:
Associate the first term with compression and the second term with retention.Ignoring the negative sign before βI(C;Y).
The word-related term appears with a negative sign, so retaining high I(C;Y) supports a lower combined objective.
Fix:
Read the complete expression, including its sign, before deciding what minimization favors.Calling the expression only a description of relationships.
It is a clustering objective because the method searches over probabilistic point-to-cluster assignments and evaluates each one.
Fix:
Include the candidate assignments and the minimization step in the explanation.
When reading an information bottleneck clustering objective, use this order: identify the variables in each mutual information term, determine whether the term is being reduced or retained, then connect the expression to the probabilistic assignments being searched.
Check Your Understanding
Explain in your own words why I(X;C) − βI(C;Y) can serve as a clustering objective for documents.
Hints
- Start by identifying what X, C, and Y represent.
- Explain the compression role of I(X;C).
- Explain the retention role of I(C;Y) and the effect of its negative sign.
- Finish by describing the search over probabilistic point-to-cluster assignments.
What do you think happens?
Before reading the explanation, predict which part of the objective represents compression and which part represents retention of word-related information.
Reveal answer
Answer: I(X;C) represents compression; I(C;Y) represents retained word-related information.
The first term measures the document–cluster connection and is sought to be small. The second measures the cluster–word connection and is retained with a negative weighted sign in the objective.
Key Takeaways
- Mutual information measures dependence between two random variables.
- In document clustering, X is document identity, C is the clustering variable, and Y is word identity.
- I(X;C) is associated with compression of document information, while high I(C;Y) represents retention of relevant word-related information.
- β weights the word-related term and participates in the compression–retention tradeoff.
- The expression is a clustering objective because it is evaluated while searching over probabilistic point-to-cluster assignments.
Key Takeaways
- Mutual information I(X;Y) measures dependence between two random variables.
- The information bottleneck objective is I(X;C) − βI(C;Y).
- Small I(X;C) compresses document-specific information, while high I(C;Y) retains information associated with words.
- X, C, and Y identify documents, clusters, and words; β weights the retention term.
- Minimizing the objective over probabilistic assignments makes the expression a clustering objective.