Concepts / Machine Learning Essentials

Machine Learning Essentials

Unsupervised learning finds useful transformations of input data without target values.

  • Programming

A Dataset Without Answers

Imagine receiving a dataset containing observations and measurements, but no column that states the correct answer for each observation. You can still investigate the dataset. You may search for a more useful representation, make the data easier to view, reduce its size, remove noise, or examine relationships among its measurements. This setting is called unsupervised learning.

paired withInput datawith target valuesInput datawithout target valuesTarget valuescorrect answers
What enters an unsupervised-learning method, and what is missing compared with supervised learning?

Unsupervised learning finds useful transformations of input data without target values.

What do you think happens?

A dataset contains observations but no column identifying the correct answer for each observation. Is this a setting in which unsupervised learning can be used?

  • Yes
  • No
Reveal answer

Answer: Yes

Unsupervised learning operates on input data without target values. It can investigate the data by seeking useful representations, groups, or relationships.

From Observations to Representation

The word transformation is central to this topic. Unsupervised learning does not require a target value telling the method what the correct output should be. Instead, it works with the input observations and seeks a representation that is useful for a particular analytical purpose.

provided toproducesInput observationsno target valuesUnsupervised methodfinds a usefultransformationUseful representationsupports analysis
How does raw input data move through an unsupervised method to become a more useful representation?

Choosing the Right Unsupervised Goal

Suppose a dataset records many measurements for each observation. You want to make the data easier to view and store, without using target values. What kind of unsupervised-learning goal fits this need?

Identify the available information: The dataset contains observations and measurements, but no target values.

Identify the desired change: The goal is a representation with fewer dimensions so the data is easier to visualize or store.

Choose the category: Seeking a representation with fewer dimensions is dimensionality reduction.

The task belongs to dimensionality reduction, a category of unsupervised learning.

Two Different Analytical Questions

Dimensionality reduction and clustering both begin with input data that has no targets, but they answer different questions. Dimensionality reduction asks how to represent the observations using fewer dimensions. Clustering asks which observations appear to belong together.

represented with fewer dimensionsorganized by similarityInput datamany measurementsInput dataobservationsFewer dimensionssmaller representationGroupssimilar data points
How does dimensionality reduction change the representation of data, and how does clustering instead group similar data points?
CategoryMain questionResult
Dimensionality reductionCan the observations be represented with fewer dimensions?A representation with fewer dimensions
ClusteringWhich observations appear to belong together?Groups among data points

If the main problem is that each observation has many measurements and you need a smaller representation, dimensionality reduction fits the question. If the main problem is discovering which observations seem to belong together, clustering fits the question.

Why Explore Before Prediction

Unsupervised learning can be useful before attempting a supervised-learning problem because it helps analysts investigate the dataset before target labels and a supervised model are introduced. The investigation may reveal useful representations, make the data easier to visualize, reduce its size, remove noise, or expose correlations among measurements.

investigatecan revealcan informInput datasetobservations withouttargetsUnsupervised analysisinvestigate the dataUseful representationpatterns or relationshipsSupervised problemintroduced afterward
What can unsupervised analysis reveal or improve before target labels and a supervised model are introduced?
PurposeHow the data may become more useful
VisualizationA representation can make the data easier to view.
CompressionA representation can reduce the size of the data.
DenoisingThe data can be investigated for the purpose of removing noise.
Understanding correlationsRelationships among measurements can be examined.

Practical purposes of unsupervised learning identified in the source material.

Mistakes in Choosing a Category

  • Treating the absence of targets as a reason that no analysis is possible.

    Unsupervised learning is specifically designed to find useful transformations of input data without target values.

    Fix: Ask what representation, grouping, or relationship would make the input data more useful.

  • Calling every unsupervised task clustering.

    That goal is dimensionality reduction, not clustering.

    Fix: Use dimensionality reduction when the desired result is a representation with fewer dimensions.

  • Calling every task that changes representation dimensionality reduction.

    Grouping data points according to similarity is the intuition behind clustering.

    Fix: Use clustering when the desired result is groups among data points.

  • Waiting until after supervised learning to investigate the dataset.

    Unsupervised learning can be a practical way to investigate a dataset before attempting a supervised-learning problem.

    Fix: Consider exploratory unsupervised analysis before introducing target labels and a supervised model.

Practice and Summary

EASY

A dataset contains many measurements for each observation and no target values. First, you want a smaller representation for easier viewing. Next, you want to discover which observations appear to belong together. Name the unsupervised-learning category for each goal.

Hints
  • Focus on the desired result of the first task.
  • A smaller representation with fewer dimensions is different from groups among data points.
  • The second task asks which observations appear to belong together.

Practice Solution

Match each goal to dimensionality reduction or clustering.

First goal: A smaller representation for easier viewing means seeking fewer dimensions, so this is dimensionality reduction.

Second goal: Discovering which observations appear to belong together means seeking groups among data points, so this is clustering.

The first goal is dimensionality reduction. The second goal is clustering.

  1. Unsupervised learning works with input data without target values. It seeks useful transformations or ways to investigate observations. Dimensionality reduction seeks a representation with fewer dimensions, while clustering seeks groups among data points. These methods can support visualization, compression, denoising, and understanding correlations, including during exploration before a supervised-learning problem.

Key Takeaways

  • Unsupervised learning finds useful transformations of input data without target values.
  • It can help investigate a dataset before attempting a supervised-learning problem.
  • Dimensionality reduction seeks a representation with fewer dimensions.
  • Clustering seeks groups among data points.
  • Practical purposes include visualization, compression, denoising, and understanding correlations.