Data representation
Machine learning discovers data-processing rules from examples of inputs and expected outputs.
From Examples to Rules
A machine-learning system is not given a complete list of hand-written rules for every possible input. Instead, it is exposed to examples of inputs and expected outputs. From those examples, it discovers data-processing rules for carrying out a task.
Data representation is central to this process. A representation is a way to encode or view data. The system does not only need to produce an output; it also needs to find a way of viewing the input that makes producing the expected output easier.
The Learning Feedback Loop
Learning requires more than input data alone. The system needs examples of inputs, expected outputs for those examples, a way to measure performance, and feedback that is used to adjust the algorithm. The examples provide material for discovering rules. The expected outputs provide a target for the task. The performance measurement indicates how well the current rules are doing, and feedback guides adjustment of those rules.
A Simplified Learning Cycle
Imagine a machine-learning system that must assign an input to an expected category.
1. Provide examples: The system receives input examples together with the expected output for each example.
2. Apply current rules: The system uses its current data-processing rules to produce an output.
3. Measure performance: The produced output is considered in relation to the expected output, creating a performance measurement.
4. Use feedback: Feedback from the measurement is used to adjust the algorithm and its data-processing rules.
5. Seek a more useful view: The system can move toward a representation in which producing the expected output is easier.
The learning process uses examples, measurement, and feedback to discover rules and representations that support the task.
Why Representation Matters
The same data can be represented in different ways. Those representations do not necessarily support a task equally well. A task that is difficult with one representation may become easier with another because the useful information is presented in a form that is more directly usable for the task.
Generated example: Suppose a system receives descriptions of items and must assign each item to an expected category. One representation might preserve the descriptions in a form that does not make the category distinction direct. Another representation might organize the available information so that the distinction is easier for the task to use. The input data has not necessarily changed; the way it is encoded or viewed has changed.
Representation Learning and Prediction
Learning representations means finding transformations of input data that move the model closer to the expected output. The model's task is to discover rules that turn available input into a form useful for the task, such as classification. This search for appropriate representations is therefore connected directly to the central goal of machine learning: discovering data-processing rules from examples of inputs and expected outputs.
Data representation is a way to encode or view data. In machine learning, representation learning is the search for transformations of input data that produce a form useful for carrying out the task and moving toward the expected output.
| Question | Representation-focused answer | Why it matters |
|---|---|---|
| What is being changed? | The way input data is encoded or viewed | Different views can support different tasks |
| What guides the change? | Examples, expected outputs, performance measurement, and feedback | The system can adjust its data-processing rules |
| What is the goal? | A representation useful for the task | Producing the expected output becomes easier |
Mistakes About Representation
Treating representation as unrelated to the task
Different representations of the same data can support different tasks more effectively.
Fix:
Ask whether the representation makes the information needed for the expected output more directly usable.Describing learning as memorizing one answer per example
The system is described as discovering data-processing rules from examples, not merely keeping a list of answers.
Fix:
Focus on the rules and transformations that turn available input into a form useful for the task.Leaving feedback out of the learning process
Learning requires performance measurement and feedback used to adjust the algorithm.
Fix:
Trace the full cycle: examples and expected outputs, produced output, performance measurement, feedback, and adjusted rules.Assuming that the output alone explains learning
The important question also includes how the system changes the way it views the input so that producing the expected output becomes easier.
Fix:
Examine both the produced output and the representation used to reach it.
When analyzing a machine-learning process, name both the task and the representation. Then ask how examples, performance measurement, and feedback connect the current representation and rules to the expected output.
Check Your Understanding
A machine-learning system receives input examples and expected outputs. Its current rules produce outputs, and a performance measurement generates feedback that adjusts the algorithm. Explain where data representation appears in this process and why changing the representation could make the task easier.
Hints
- Start by defining representation as a way to encode or view data.
- Identify the transformation from input data to a form useful for the task.
- Connect the useful form to the expected output and the feedback used for adjustment.
Tracing the Ingredients
Identify the role of each ingredient in a learning process: input examples, expected outputs, performance measurement, feedback, and representation.
Input examples: These give the system data from which it can discover data-processing rules.
Expected outputs: These provide the result toward which the task is directed.
Performance measurement: This indicates how the current produced output relates to the expected output.
Feedback: This is used to adjust the algorithm and its rules.
Representation: This is the way the input is encoded or viewed; a useful representation can make producing the expected output easier.
All five ingredients fit together: examples and expected outputs define the learning situation, measurement and feedback guide adjustment, and representation determines how directly the input supports the task.
Key Takeaways
- A data representation is a way to encode or view data.
- Machine learning discovers data-processing rules from examples of inputs and expected outputs rather than beginning with a complete hand-written rule for every input.
- Learning requires input data, performance measurement, and feedback used to adjust the algorithm.
- Different representations of the same data can support different tasks more effectively.
- Representation learning searches for transformations that make the input more useful for producing the expected output, making it central to machine learning and deep learning.
Key Takeaways
- Data representation describes how data is encoded or viewed.
- Machine learning uses examples and expected outputs to discover data-processing rules.
- Performance measurement and feedback guide adjustments to the algorithm.
- A useful representation can make a task easier without reducing learning to memorizing answers.
- Representation learning connects transformations of input data with the central goal of producing useful predictions or decisions.