Concepts / Inference-Time Prediction

Inference-Time Prediction

Model ensembling pools outputs from multiple models rather than relying on one model.

  • Programming

One Input, Several Opinions

Suppose several models have been trained for the same task. At inference time, you have a choice: use only one model's output, or bring the outputs from several models together to form the final prediction. Model ensembling is the second strategy. It pools outputs from multiple models rather than relying on only one.

evaluateevaluateevaluateproduceproduceproducecombinecombinecombineformInputsame data itemModel Atrained modelPrediction Amodel outputPrediction Poolcombined outputsFinal Predictionensemble outputModel Btrained modelPrediction Bmodel outputModel Ctrained modelPrediction Cmodel output
How does one input move through several trained models and become one ensemble prediction?

Tracing the Model Outputs

The important change occurs at inference time. The input is presented to multiple already-trained models. Each model produces an output for that input. Instead of selecting one output immediately, the ensemble pools the collection and uses it to form the final prediction.

Pooling Three Classifier Predictions

Three classifiers predict whether the same input belongs to Class A or Class B. Their Class A scores are 0.60, 0.80, and 0.70. Use a straightforward averaging method to pool these predictions.

Collect: Keep the Class A prediction from each classifier: 0.60, 0.80, and 0.70.

Combine: Add the three Class A scores: 0.60 + 0.80 + 0.70 = 2.10.

Average: Divide the combined score by the three contributing classifiers: 2.10 divided by 3 equals 0.70.

Interpret: The ensemble's pooled Class A score for this illustrative example is 0.70.

Averaging the classifier predictions produces an ensemble Class A score of 0.70.

The averaging step does not choose one model as the winner before combining the outputs. It uses the predictions from all of the participating classifiers. This is what makes the collection an active ensemble rather than merely a set of models stored side by side.

Averaging Classifier Predictions

A straightforward pooling method is to average classifier predictions at inference time. For each class, the predictions supplied by the participating classifiers are brought together and averaged to produce the ensemble prediction for that class.

poolpoolpoolproduceClassifier AClass A: 0.60Average0.70Ensemble PredictionClass A: 0.70Classifier BClass A: 0.80Classifier CClass A: 0.70
How are class-probability predictions from several classifiers averaged into one ensemble prediction?

Perspectives That Add Value

Independent models may be useful for different reasons and may attend to different aspects of the data. Their predictions can therefore contribute different perspectives to the combined view. The benefit of ensembling comes from combining these perspectives, not simply from counting how many models have been trained.

If every model contributed exactly the same perspective, pooling their outputs would add less information. When evaluating a proposed ensemble, ask two questions: Are several models available, and what does each model contribute to the final combined view?

select onepool outputsSeveral Modelsmodels existOne Model Outputselected outputActive Ensembleoutputs are pooledCombined Predictionensemble output
What is the difference between having multiple trained models available and using their outputs together at inference time?

Mistakes in Ensemble Reasoning

  • Calling a collection of trained models an ensemble even when only one model's output is used.

    The models exist, but their outputs are not being pooled to form the prediction.

    Fix: Use the term ensemble when multiple model outputs are brought together for the final prediction.

  • Assuming that more models automatically means a more useful ensemble.

    The source identifies combined perspectives, rather than model count alone, as the important benefit.

    Fix: Consider what distinct useful perspective each model contributes.

  • Describing averaging as choosing the strongest single model.

    Averaging pools the classifier predictions instead of selecting one model's output.

    Fix: Bring the participating classifier predictions together and average them at inference time.

Reasoning Check

EASY

A team has trained four models for the same task. During inference, it runs all four models, collects their classifier predictions, and averages those predictions before producing the final output. Is this an ensemble? Explain what makes it one, and identify the pooling method used in this example.

Hints
  • Ask whether more than one model output contributes to the final prediction.
  • Look for the operation applied to the classifier predictions.
  • Separate the existence of several models from their use at inference time.

What do you think happens?

A system has three trained models, but it uses only Model A's output and discards the other two. Is the system using an ensemble?

  • Yes, because several models have been trained
  • No, because only one model's output is used
  • Yes, because the unused models still exist
Reveal answer

Answer: No, because only one model's output is used.

Having several trained models is not enough. An ensemble uses the collection of model outputs to form the final prediction.

Final Takeaways

  1. Model ensembling combines predictions from multiple models instead of relying on one model.
  2. At inference time, the same input can be passed through several trained models before their outputs are pooled.
  3. Averaging classifier predictions is a straightforward method for producing an ensemble prediction.
  4. Independent models may contribute different useful perspectives on the data.
  5. Several trained models become an active ensemble only when their outputs are used together to form the final prediction.

Key Takeaways

  • Model ensembling pools outputs from multiple models at inference time.
  • Averaging is a straightforward way to combine classifier predictions.
  • The value of an ensemble depends on useful differences in the perspectives contributed by its models.
  • Having several trained models is not the same as using an ensemble; their outputs must contribute to the final prediction.