Machine Learning Essentials
Unsupervised learning finds useful transformations of input data without target values.
A Dataset Without Answers
Imagine receiving a dataset containing observations and measurements, but no column that states the correct answer for each observation. You can still investigate the dataset. You may search for a more useful representation, make the data easier to view, reduce its size, remove noise, or examine relationships among its measurements. This setting is called unsupervised learning.
Unsupervised learning finds useful transformations of input data without target values.
What do you think happens?
A dataset contains observations but no column identifying the correct answer for each observation. Is this a setting in which unsupervised learning can be used?
Reveal answer
Answer: Yes
Unsupervised learning operates on input data without target values. It can investigate the data by seeking useful representations, groups, or relationships.
From Observations to Representation
The word transformation is central to this topic. Unsupervised learning does not require a target value telling the method what the correct output should be. Instead, it works with the input observations and seeks a representation that is useful for a particular analytical purpose.
Choosing the Right Unsupervised Goal
Suppose a dataset records many measurements for each observation. You want to make the data easier to view and store, without using target values. What kind of unsupervised-learning goal fits this need?
Identify the available information: The dataset contains observations and measurements, but no target values.
Identify the desired change: The goal is a representation with fewer dimensions so the data is easier to visualize or store.
Choose the category: Seeking a representation with fewer dimensions is dimensionality reduction.
The task belongs to dimensionality reduction, a category of unsupervised learning.
Two Different Analytical Questions
Dimensionality reduction and clustering both begin with input data that has no targets, but they answer different questions. Dimensionality reduction asks how to represent the observations using fewer dimensions. Clustering asks which observations appear to belong together.
| Category | Main question | Result |
|---|---|---|
| Dimensionality reduction | Can the observations be represented with fewer dimensions? | A representation with fewer dimensions |
| Clustering | Which observations appear to belong together? | Groups among data points |
If the main problem is that each observation has many measurements and you need a smaller representation, dimensionality reduction fits the question. If the main problem is discovering which observations seem to belong together, clustering fits the question.
Why Explore Before Prediction
Unsupervised learning can be useful before attempting a supervised-learning problem because it helps analysts investigate the dataset before target labels and a supervised model are introduced. The investigation may reveal useful representations, make the data easier to visualize, reduce its size, remove noise, or expose correlations among measurements.
| Purpose | How the data may become more useful |
|---|---|
| Visualization | A representation can make the data easier to view. |
| Compression | A representation can reduce the size of the data. |
| Denoising | The data can be investigated for the purpose of removing noise. |
| Understanding correlations | Relationships among measurements can be examined. |
Practical purposes of unsupervised learning identified in the source material.
Mistakes in Choosing a Category
Treating the absence of targets as a reason that no analysis is possible.
Unsupervised learning is specifically designed to find useful transformations of input data without target values.
Fix:
Ask what representation, grouping, or relationship would make the input data more useful.Calling every unsupervised task clustering.
That goal is dimensionality reduction, not clustering.
Fix:
Use dimensionality reduction when the desired result is a representation with fewer dimensions.Calling every task that changes representation dimensionality reduction.
Grouping data points according to similarity is the intuition behind clustering.
Fix:
Use clustering when the desired result is groups among data points.Waiting until after supervised learning to investigate the dataset.
Unsupervised learning can be a practical way to investigate a dataset before attempting a supervised-learning problem.
Fix:
Consider exploratory unsupervised analysis before introducing target labels and a supervised model.
Practice and Summary
A dataset contains many measurements for each observation and no target values. First, you want a smaller representation for easier viewing. Next, you want to discover which observations appear to belong together. Name the unsupervised-learning category for each goal.
Hints
- Focus on the desired result of the first task.
- A smaller representation with fewer dimensions is different from groups among data points.
- The second task asks which observations appear to belong together.
Practice Solution
Match each goal to dimensionality reduction or clustering.
First goal: A smaller representation for easier viewing means seeking fewer dimensions, so this is dimensionality reduction.
Second goal: Discovering which observations appear to belong together means seeking groups among data points, so this is clustering.
The first goal is dimensionality reduction. The second goal is clustering.
- Unsupervised learning works with input data without target values. It seeks useful transformations or ways to investigate observations. Dimensionality reduction seeks a representation with fewer dimensions, while clustering seeks groups among data points. These methods can support visualization, compression, denoising, and understanding correlations, including during exploration before a supervised-learning problem.
Key Takeaways
- Unsupervised learning finds useful transformations of input data without target values.
- It can help investigate a dataset before attempting a supervised-learning problem.
- Dimensionality reduction seeks a representation with fewer dimensions.
- Clustering seeks groups among data points.
- Practical purposes include visualization, compression, denoising, and understanding correlations.