Data Compression
Unsupervised learning finds useful transformations of input data without target values.
When Answers Are Missing
Imagine receiving a dataset containing observations but no column that states the correct answer for each observation. You cannot train by comparing predictions with supplied target values, but the dataset can still reveal useful structure. Unsupervised learning investigates input data without target values and can produce useful transformations of that data.
The defining feature is not that the data is empty of information. It is that the input observations arrive without target values.
From Many Measurements to Fewer
A dataset may record many measurements for every observation. One useful transformation is to seek a representation with fewer dimensions. This smaller representation can make the data easier to visualize or store. In this sense, data compression is one practical purpose of finding useful transformations in unsupervised learning.
Two Different Unsupervised Questions
Dimensionality reduction and clustering are both well-known categories of unsupervised learning, but they answer different questions. Dimensionality reduction asks how to represent observations with fewer dimensions. Clustering asks which observations appear to belong together.
| Category | Main question | Result sought |
|---|---|---|
| Dimensionality reduction | Can the observations be represented with fewer dimensions? | A representation with fewer dimensions |
| Clustering | Which observations appear to belong together? | Groups of similar data points |
A Compression Decision
Choosing the Relevant Unsupervised Task
An analyst has observations containing many measurements. The analyst wants a smaller representation that is easier to visualize or store, but does not have target values.
Identify the available data: The dataset contains input observations, but no target values. That places the task in the setting of unsupervised learning.
Identify the requested change: The goal is a representation with fewer dimensions, not a set of groups of similar observations.
Select the category: Because the goal is a smaller representation, the task belongs to dimensionality reduction.
Connect the purpose: The smaller representation can support visualization or storage, which are practical purposes of unsupervised learning.
This is an unsupervised dimensionality-reduction task used for data compression and easier visualization or storage.
The key reasoning step is to separate the input condition from the task goal. The absence of target values identifies the broader learning setting. The requested outcome then distinguishes compression-oriented dimensionality reduction from clustering.
Uses Before Supervised Learning
Unsupervised learning can be useful before attempting a supervised-learning problem because it helps analysts investigate the dataset first. A learner can help with visualization, compression, and denoising, and can help analysts understand correlations among measurements. These activities can reveal useful properties of the input data before target-based learning is attempted.
- Visualization: create a representation that makes the data easier to view.
- Compression: seek a smaller representation of the input data.
- Denoising: support the removal of noise from the data.
- Understanding correlations: examine relationships among measurements.
Common Classification Mistakes
Treating unsupervised learning as learning with no input data.
Unsupervised learning begins with input data. What is absent is the target value for each observation.
Fix:
Describe the setting as input observations without target values.Using dimensionality reduction and clustering as if they had the same goal.
Dimensionality reduction seeks a representation with fewer dimensions, while clustering seeks groups among data points.
Fix:
Ask whether the desired result is a smaller representation or groups of similar observations.Defining compression only as grouping observations.
Grouping observations is the intuition behind clustering. Compression is associated with seeking a smaller representation.
Fix:
For a smaller representation, connect the task to dimensionality reduction and compression.Waiting until after supervised learning to investigate the dataset.
Unsupervised learning can support investigation before attempting a supervised-learning problem.
Fix:
Consider whether an unsupervised transformation could clarify the input data first.
Practice Check
A dataset contains observations with many measurements and no target values. An analyst wants to discover which observations appear to belong together. Is this primarily dimensionality reduction or clustering? Explain why.
Hints
- Start by identifying whether target values are present.
- Then focus on whether the requested result is fewer dimensions or groups of observations.
What do you think happens?
Which category best matches the practice task?
Reveal answer
Answer: Clustering
The task asks which observations appear to belong together. Grouping data points according to similarity is the intuition behind clustering.
Key Takeaways
- Unsupervised learning finds useful transformations of input data without target values.
- Data compression can be approached as seeking a smaller representation of observations.
- Dimensionality reduction seeks fewer dimensions, while clustering seeks groups among data points.
- Unsupervised learning can support visualization, compression, denoising, and understanding correlations.
- These uses make unsupervised learning a practical way to investigate data before attempting a supervised-learning problem.
Key Takeaways
- Unsupervised learning works with input observations but no target values.
- Compression is a practical purpose of transforming data into a smaller representation.
- Dimensionality reduction and clustering are different unsupervised-learning categories.
- Unsupervised analysis can support visualization, denoising, and understanding correlations before supervised learning.