Concepts / Fighting Overfitting

Fighting Overfitting

A machine learning problem is defined by its inputs, predicted outputs, and problem type.

  • Programming

Before Choosing a Model

Overfitting is often discussed as a model problem, but a deeper problem can appear earlier: the machine learning task may not be defined in a way that makes reliable prediction possible. Before choosing a model, start by identifying what information is available, what the model should predict, and what kind of prediction task you are solving.

A machine learning problem is defined by its inputs, predicted outputs, and problem type.

Three Parts of a Task

A useful problem definition connects three parts. Inputs are the information supplied to the model. Predicted outputs are the results the model is expected to produce. The problem type describes the structure of that output and the task the model must perform. These parts are connected: the inputs must be considered in relation to the desired outputs, and the problem type guides choices such as model architecture and loss function.

informdefineguidesInputsinformation suppliedModel choicesarchitecture and lossPredicted outputsresults expectedProblem typetask structure
How are the available inputs, predicted outputs, and problem type connected when defining a machine learning task?

Separating the Task Parts

Suppose a model receives available information and must produce a prediction. Define the task without choosing a model yet.

Identify the inputs: Write down the information that will be supplied to the model.

Identify the output: State exactly what the model is expected to predict.

Identify the problem type: Describe the structure of the prediction, such as a class, a scalar, a vector, or another listed task type.

Delay model selection: Only after the task is defined should the problem type guide choices such as model architecture and loss function.

A clear task definition states the inputs, the predicted outputs, and the problem type before a model is selected.

Data as a Limitation

Data availability is often the first practical limitation on a machine learning problem. To learn from examples, you need input examples and known target outputs. However, having input-target examples does not prove that the inputs contain enough information for prediction. It also does not prove that the available data represents the relationship you want to learn.

must containhelp definesupportssupportsInput examplesavailable informationUseful informationinputs relate to outputsMachine learning taskcan be attemptedKnown targetsdesired outputsRepresentative datarelationship to learn
How does the availability and usefulness of input-target examples determine whether a machine learning problem can be attempted?

The Two Working Hypotheses

Defining a machine learning problem involves two important hypotheses. First, the inputs contain useful information about the outputs; otherwise, the model has no informative basis for making the prediction. Second, the available data represents the relationship that you want to learn, including the behavior that will matter when the model is used. These are hypotheses because examples alone do not automatically establish that either condition is true.

supports checkingsupports checkingrequired forrequired forAvailable datainput-target examplesUseful input signalinputs inform outputsDefined problemprediction is worthattemptingRepresentativerelationshipdata reflects what to learn
What two assumptions connect available training data to a meaningful machine learning problem?

When Patterns Move

A problem can be nonstationary when the relationship between inputs and outputs changes over time. A model trained on earlier data may then make increasingly incorrect predictions because the pattern it learned no longer matches the current relationship. Nonstationary problems require attention to the changing relationship and to the relevant time scales.

model learnsmakes past pattern less reliableapplies old relationshipPast relationshipinputs linked to earlieroutputsTrained modellearned from past dataChangedrelationshipcurrent inputs linkdifferentlyCurrent predictionsincreasingly incorrect
How does a change in the relationship between inputs and outputs cause a model trained on past data to become increasingly incorrect?

When time matters, do not ask only whether the training examples are available. Ask whether the relationship between inputs and outputs is stable at the time scale relevant to use. If that relationship changes, treat the task as potentially nonstationary.

Why the Future Must Resemble the Past

Machine learning uses patterns found in past training data to make predictions about future behavior. This only works when the future retains enough resemblance to the patterns represented in the training data. If the future no longer resembles the past, the model's learned relationship may no longer provide useful predictions. This is the central limitation behind using historical examples to predict what comes next.

revealis applied toresembles pastno longer resembles pastPast examplestraining dataLearned patternrelationship from examplesFuture behaviorresembles past or changesUseful predictionwhen patterns remainsimilarPoor predictionwhen patterns change
How does a model use patterns in past training data to predict future behavior, and what happens when the future no longer resembles the past?

What do you think happens?

A model was trained on examples from an earlier period. If the relationship between its inputs and outputs changes later, what should you expect?

  • The old model is guaranteed to remain equally accurate
  • The old model may make increasingly incorrect predictions
  • The model automatically learns the new relationship without new information
Reveal answer

Answer: The old model may make increasingly incorrect predictions.

The model learned a relationship from past data. When that relationship changes, the past pattern may no longer match current behavior.

Task Types and Model Choices

Problem typePredicted output structureWhy it matters
Binary classificationOne of two classesThe task type helps guide model architecture and loss function.
Multiclass classificationOne class among multiple classesThe task type helps guide model architecture and loss function.
Scalar regressionA scalar valueThe task type helps guide model architecture and loss function.
Vector regressionA vector of valuesThe task type helps guide model architecture and loss function.
Multiclass multilabel classificationMultiple class labelsThe task type helps guide model architecture and loss function.
ClusteringGroups discovered in dataThe task type identifies a different machine learning task structure.
GenerationGenerated outputThe task type identifies a different machine learning task structure.
Reinforcement learningA learning task based on interactionThe task type identifies a different machine learning task structure.

Do not treat the problem type as a label added after modeling. It is part of the definition of the machine learning problem. Whether the task is binary classification, multiclass classification, scalar regression, vector regression, multiclass multilabel classification, clustering, generation, or reinforcement learning affects how the task is understood and guides choices such as model architecture and loss function.

Mistakes in Problem Definition

  • Assuming that input-target pairs automatically make the task solvable.

    The inputs may not contain useful information about the outputs, and the available data may not represent the relationship that must be learned.

    Fix: Check both whether the inputs carry useful information and whether the data represents the relationship of interest.

  • Ignoring the problem type.

    Problem type guides choices such as model architecture and loss function.

    Fix: Name the problem type as part of the initial task definition.

  • Treating historical patterns as permanently stable.

    A changing relationship can make the problem nonstationary and cause increasingly incorrect predictions.

    Fix: Consider changing relationships and the relevant time scales before relying on past training data.

Apply the Definition

MEDIUM

Describe a hypothetical machine learning task in four parts: the inputs, the predicted outputs, the problem type, and the assumption that must hold for past training data to help with future predictions. Then state one reason the task might fail even if input-target examples are available.

Hints
  • Start by naming the information supplied to the model.
  • Describe the output structure before naming the problem type.
  • Check whether the inputs contain useful information about the outputs.
  • Check whether the relationship represented by the data is likely to remain relevant in the future.
  1. A strong machine learning problem definition identifies the inputs, predicted outputs, and problem type. Data availability is only the starting point: the inputs must contain useful information, and the available data must represent the relationship to be learned. Machine learning relies on the future resembling patterns found in past training data. When the input-output relationship changes over time, the problem may be nonstationary and predictions may become increasingly incorrect.

Key Takeaways

  • Define a machine learning problem by its inputs, predicted outputs, and problem type.
  • Input-target examples do not guarantee that the inputs contain useful information or that the data represents the relationship to be learned.
  • The two key hypotheses are that inputs contain useful information about outputs and that available data represents the relationship of interest.
  • Future predictions rely on the future resembling patterns in past training data.
  • Changing input-output relationships can create nonstationary problems and make predictions increasingly incorrect.