Concepts / Information Theory for Machine Learning

Information Theory for Machine Learning

Mutual information I(X;Y) measures dependence between two random variables.

  • Programming

From Documents to Clusters

A clustering method replaces many individual data points with a smaller description: a cluster identity. That replacement is useful only when the clusters discard unnecessary detail while preserving information that matters for the task. Information theory provides a way to describe this tension using mutual information.

Mutual information, written I(X;Y), measures dependence between two random variables. In this setting, the question is not simply whether documents and words exist, but how much knowing one variable is connected to knowing the other. The information bottleneck approach uses this idea to decide what a compressed representation should preserve.

increasing dependenceincreasing dependenceIndependent variableslow shared informationWeakly relatedvariablessome shared informationStrongly dependentvariableshigh shared information
How does the amount of shared information change as the dependence between two random variables changes?

The Four Symbols

For document clustering, X represents the identity of a document, C represents the clustering variable, and Y represents the identity of words. The clustering variable takes values in the available cluster labels. β is a weight in the information bottleneck objective.

SymbolMeaningRole in the objective
XIdentity of a documentThe original document information being compressed
CClustering variableThe compressed representation, taking values in available cluster labels
YIdentity of wordsThe information treated as relevant to the document
βWeightControls the weight applied to the retained-information term

The variables and parameter used in the document-clustering formulation.

combined withfollowed byweightsI(X;C)document–cluster connection−retention term issubtractedβweightI(C;Y)cluster–word connection
What does each symbol represent, and how does β participate in balancing compression with retained word-related information?

Balancing Compression and Relevance

I(X;C) − βI(C;Y)

The first term, I(X;C), measures the connection between a document's identity and its cluster identity. Making this quantity small expresses compression: the cluster representation should not preserve all of the original document-specific detail.

The second term, I(C;Y), measures the connection between the clustering variable and the identity of words. It appears with a negative sign and a weight β. The stated goal is to retain high mutual information here because words are treated as the relevant information about the document.

compresses intodoes not preserve allpreserves connection toDocument identity Xoriginal document detailCluster variable Csmaller descriptionDocument-specificdetailinformation not preservedWord identity Yrelevant informationretained
As document identity X is compressed into cluster variable C, what information is discarded and what word-related information is preserved?

Interpreting a Candidate Clustering

Consider a candidate clustering represented by C. How should its two mutual-information terms be interpreted?

Inspect I(X;C): This term asks how much the cluster identity is connected to the identity of each document. A small value means the representation has compressed away more document-specific detail.

Inspect I(C;Y): This term asks how much the cluster identity is connected to word identity. A high value means the representation retains more of the word-related information treated as relevant.

Combine the terms: The objective combines the compression term with the negatively weighted retention term: I(X;C) − βI(C;Y).

A useful candidate has a compressed cluster representation while retaining information about words. The objective expresses both requirements together.

Why Assignments Create Clusters

The expression becomes a clustering objective because the method searches over probabilistic point-to-cluster assignments. Each assignment describes how documents are associated with the available values of C. The objective then evaluates that assignment through the combined compression and retained-information terms.

provide points to assigndefinesis evaluated bymeasuresDocument datapointsdocument identities XProbabilisticassignmentdocuments to cluster labelsCandidate clusteringclustering variable CInformationbottleneck objectiveI(X;C) − βI(C;Y)Assignment evaluationcompression and wordinformation
How do different probabilistic assignments of documents to cluster representations become candidate clusterings evaluated by the objective?

Changing the probabilistic assignments changes the relationship between X and C, and it can also change the relationship between C and Y. Therefore, the method is not merely describing two fixed relationships. It is comparing possible assignments by how well they combine compression of document information with retention of word-related information.

When reading this objective, first name the variables in each mutual-information term, then ask what that pairing is intended to accomplish. This prevents the common mistake of treating the formula as one undifferentiated expression.

Common Interpretation Errors

  • Treating C as another document variable

    C is the clustering variable: the smaller description that takes values in available cluster labels.

    Fix: Read X as document identity and C as cluster identity.

  • Assuming compression means preserving every document detail

    Small I(X;C) expresses the goal of not preserving all original document-specific detail.

    Fix: Interpret compression as discarding unnecessary document detail while retaining information relevant to the task.

  • Ignoring the negative sign before βI(C;Y)

    The second term is negatively weighted because the stated goal is to retain high mutual information between C and Y.

    Fix: Separate the goals: reduce I(X;C), while retaining high I(C;Y).

  • Calling the expression a clustering objective without mentioning assignments

    The clustering interpretation comes from searching over probabilistic point-to-cluster assignments and evaluating each one.

    Fix: Explain both the formula and the assignments that the method evaluates.

Check Your Interpretation

MEDIUM

For the objective I(X;C) − βI(C;Y), write one sentence explaining the purpose of each term. Then explain why changing the probabilistic document-to-cluster assignments changes the clustering evaluation.

Hints
  • Start by identifying what each pair of variables represents.
  • Use compression for the first term and retained word-related information for the second term.
  • Mention that the method evaluates possible assignments rather than one fixed assignment.

What do you think happens?

If a candidate representation has small I(X;C) but does not retain much information about Y, does it satisfy the full information bottleneck goal?

  • Yes, because compression is the only goal
  • No, because the representation should also retain relevant word-related information
  • Yes, because I(C;Y) is unrelated to clustering
Reveal answer

Answer: No, because the representation should also retain relevant word-related information.

The objective balances small I(X;C), which expresses compression, with high I(C;Y), which expresses retention of information associated with words.

Key Takeaways

  1. Mutual information I(X;Y) measures dependence between two random variables.
  2. In document clustering, X is document identity, C is the clustering variable, and Y is word identity.
  3. The objective I(X;C) − βI(C;Y) seeks compression of document information while retaining word-related information.
  4. Small I(X;C) represents compression, while high I(C;Y) represents retention of relevant information.
  5. The expression becomes a clustering objective because it evaluates probabilistic assignments of documents to cluster labels.

Key Takeaways

  • Mutual information measures dependence between random variables.
  • The information bottleneck objective balances reducing document-specific information with preserving word-related information.
  • X, C, and Y represent document identity, cluster identity, and word identity; β weights the retained-information term.
  • Searching over probabilistic document-to-cluster assignments makes the expression an objective for clustering.