Concepts / Data Visualization

Data Visualization

Unsupervised learning finds useful transformations of input data without target values.

  • Programming

A Dataset Without Answers

Imagine receiving a dataset containing observations and their measurements, but no column that tells you the correct answer for each observation. You can still investigate the data. You might search for a more useful representation, make the observations easier to view, reduce the amount of stored data, remove noise, or examine relationships among measurements. This is the setting for unsupervised learning.

Unsupervised learning finds useful transformations of input data without target values.

paired withprovided withoutInput dataobservations andmeasurementsInput dataobservations andmeasurementsTarget valuescorrect answersNo targetsno correct-answer column
What information enters an unsupervised-learning method, and how is that different from supervised learning with target values?

Tracing the Investigation

The absence of targets does not mean that the input data is useless. It changes the question. Instead of asking a learner to reproduce a supplied correct answer, you investigate the structure already present in the measurements. The result may be a new representation of the same observations or a set of groups formed from observations that appear similar.

Choosing the Investigation

A dataset records many measurements for each observation, but it contains no target values. Decide whether each goal is dimensionality reduction or clustering.

Goal 1: make the data easier to view: A representation with fewer dimensions can make data easier to visualize. This goal belongs to dimensionality reduction.

Goal 2: find observations that belong together: Discovering which observations appear to belong together is the intuition behind clustering.

Goal 3: investigate the dataset before supervised learning: Unsupervised learning can help analysts investigate a dataset before attempting a supervised-learning problem.

Dimensionality reduction changes the representation to use fewer dimensions, while clustering seeks groups among data points. Both begin with input data that has no targets.

A useful first question is: do you want a smaller representation of the observations, or do you want to discover groups among them? The first points toward dimensionality reduction; the second points toward clustering.

Two Unsupervised Paths

Dimensionality reduction and clustering are two well-known categories of unsupervised learning. They start from the same kind of input situation: observations are available, but target values are not. Their analytical goals differ, however. Dimensionality reduction seeks a representation with fewer dimensions. Clustering seeks groups among data points.

CategoryMain questionOutput idea
Dimensionality reductionCan the observations be represented with fewer dimensions?A representation with fewer dimensions
ClusteringWhich observations appear to belong together?Groups among data points

The two categories differ in what they seek from the input data.

transformproducesupportsupporthelp examineInput datamany measurements perobservationDimensionalityreductionfewer dimensionsSmallerrepresentationeasier to visualize orstoreVisualizationinspect the dataCompressionreduce its sizeCorrelationsunderstand relationships
How does high-dimensional input data get transformed into a two- or three-dimensional representation while preserving meaningful relationships?

Following the Data Transformation

Consider a generated example in which each observation has many recorded measurements. No target column identifies a correct answer. If the immediate need is to see the observations more easily, dimensionality reduction can seek a representation with fewer dimensions. The observations are still being investigated as data, but the representation used for that investigation has changed.

The same transformation can serve more than one practical purpose. A representation with fewer dimensions may make the data easier to visualize and may reduce the amount of data that needs to be stored. Unsupervised transformations can also support denoising and help analysts understand correlations present in the data.

investigatesupportsupportsupportsupportInput dataobservations andmeasurementsUseful transformationno target values requiredVisualizationmake data easier to viewCompressionreduce data sizeDenoisingremove noiseCorrelationsunderstand relationships
How can one transformation of input data support visualization, compression, denoising, or understanding correlations?

Grouping Without Target Labels

Clustering begins with a different question from dimensionality reduction. Instead of seeking a smaller representation, you want to discover which observations appear to belong together. Grouping data points according to their similarity is the intuition behind clustering.

inspectorganizeorganizeData pointsobservations withouttargetsSimilaritycompare observationsGroup Aobservations that appeartogetherGroup Bobservations that appeartogether
How are individual data points grouped into clusters based on similarities when no target values are provided?

Suppose an analyst asks which observations appear to belong together, rather than asking for a representation with fewer dimensions. Clustering addresses that grouping question. The groups are discovered from the input data; they are not supplied as target values in the dataset described here.

Mistakes in Choosing a Goal

  • Treating the absence of targets as a reason to stop analyzing the data.

    Unsupervised learning is specifically concerned with finding useful transformations of input data without target values.

    Fix: Investigate representations, groups, data size, noise, and relationships among measurements.

  • Calling every unsupervised task dimensionality reduction.

    Dimensionality reduction seeks fewer dimensions, while clustering seeks groups among data points.

    Fix: Name the goal first: smaller representation suggests dimensionality reduction; discovered groups suggest clustering.

  • Assuming that clustering produces a smaller representation.

    The two categories address different analytical needs.

    Fix: Describe clustering as grouping data points according to their similarity.

  • Ignoring unsupervised learning until after supervised learning.

    Unsupervised learning can be useful before attempting a supervised-learning problem.

    Fix: Use unsupervised investigation to explore representations, groups, compression, denoising, and correlations before the supervised task.

Practice the Distinction

EASY

A dataset contains many measurements for each observation and no target values. Match each purpose with the unsupervised-learning idea it most directly describes: creating a representation with fewer dimensions; discovering which observations belong together; making the data easier to view; reducing the data's size; removing noise; understanding relationships among measurements.

Hints
  • Look for goals involving fewer dimensions or an easier-to-view representation.
  • Look for the goal involving observations that belong together.
  • The remaining purposes describe practical benefits of investigating input data without targets.

What do you think happens?

You have observations but no target values. If your first goal is to discover which observations appear to belong together, which category best matches that goal?

  • Dimensionality reduction
  • Clustering
  • Neither category
Reveal answer

Answer: Clustering

Clustering seeks groups among data points and uses similarity as the intuition for deciding which observations appear to belong together.

Key Takeaways

  1. Unsupervised learning finds useful transformations of input data without target values.
  2. It can help investigate a dataset before attempting a supervised-learning problem.
  3. Dimensionality reduction seeks a representation with fewer dimensions.
  4. Clustering seeks groups among data points according to their similarity.
  5. Unsupervised learning can support visualization, compression, denoising, and understanding correlations.

Key Takeaways

  • Unsupervised learning analyzes input data when target values are absent.
  • Dimensionality reduction focuses on representing observations with fewer dimensions.
  • Clustering focuses on discovering groups among observations.
  • These approaches can support visualization, compression, denoising, and understanding correlations.
  • Unsupervised investigation can be useful before a supervised-learning problem.