Concepts / Online Learning Round

Online Learning Round

Online learning has no separate training phase followed by a separate prediction phase.

  • Programming

The Timing Challenge

Imagine a learner that must make a decision before knowing whether the decision is correct. The learner predicts first, receives the true answer afterward, and uses that answer to improve later decisions. This timing pattern is the central idea of online learning.

Online learning has no separate training phase followed by a separate prediction phase. Instead, prediction and learning are interwoven across a sequence of consecutive learning rounds.

The Four Events

An online learning round follows a fixed timing pattern. First, an instance is presented. Second, the learner must make a prediction about that instance. Third, the correct label is revealed. Fourth, the learner uses that label to support predictions in later rounds. The learner therefore acts before receiving feedback, not after it.

presentthen revealuse forInstanceCurrent examplePredictionLearner decidesCorrect labelAnswer revealedFuture predictionsLabel supports laterdecisions
What happens first, second, third, and fourth during one online learning round?

A Papaya Decision

A learner must decide whether a current papaya has the target taste before knowing the correct answer.

Receive the instance: The current papaya is the instance for this round.

Make a prediction: The learner predicts the papaya's label before obtaining the correct label.

Obtain the correct label: The true answer is revealed after the prediction.

Support later predictions: The labeled papaya can now contribute to the learner's future predictions.

One round combines a prediction event with a later learning event rather than placing them in separate phases.

From Test to Training

The same example has two roles at different moments in the round. Before its correct label is known, it tests what the learner currently predicts. After the correct label is obtained, that example can serve as training information for future predictions. The role changes because the learner has moved from making a prediction to receiving feedback.

evaluate current learnercorrect label revealedsupportUnlabeled exampleTest rolePredictionBefore correct labelLabeled exampleTraining roleLater predictionsSupported by feedback
How does the same example first test the learner's prediction and then update the learner as training data?

Suppose a learner predicts the label of a papaya before tasting it. At that moment, the papaya tests the learner's current prediction. Once the papaya is tasted and its correct label is available, the same papaya becomes labeled information that can support later predictions.

A Chain of Rounds

Online learning does not stop after one prediction and one label. The round ends with information that can support a later prediction, and the next round begins with another instance. Repeating this pattern creates consecutive rounds in which each round combines prediction with learning for the rounds that follow.

correct label arrivessupportscontinues tosame round patternRound 1Instance, prediction, labelLabeled informationAvailable for laterpredictionsRound 2New instance and predictionLater roundsPattern continues
How does one learning round lead into the next, and how does the learner's information change across the sequence?

The essential sequence is not training examples followed by test examples. It is instance, prediction, correct label, and support for future predictions, repeated over consecutive rounds.

Online Learning and PAC Learning

The main difference is the schedule of events. In the PAC learning model, the learner first receives a batch of training examples, learns a hypothesis from that batch, and then applies the learned hypothesis to new examples. Online learning instead processes examples through consecutive rounds, requiring a prediction before the correct label for the current instance is available.

learn fromapply topredict before feedbackthen receivesupports next roundTraining batchPAC learningCurrent instanceOnline learningLearned hypothesisAfter the batchPredictionBefore correct labelNew examplesApply hypothesisCorrect labelSupports later rounds
What is different about the timing of training, prediction, and feedback in online learning compared with PAC learning?
FeaturePAC learningOnline learning
Training timingA batch of training examples comes firstLearning information becomes available after each round's label
Prediction timingPredictions are made after learning from the batchThe learner predicts the current instance before its correct label is known
OrganizationA training stage followed by a prediction stageConsecutive rounds that combine prediction and learning
Example roleThe batch is used for training before new examples are predictedThe current example first tests the prediction and then can support future predictions

Mistakes About the Schedule

  • Treating online learning as a separate training phase followed by a separate prediction phase.

    Online learning interweaves prediction and learning across consecutive rounds.

    Fix: Place the prediction before the current correct label, then use that label to support later predictions.

  • Assuming the correct label is available before the prediction.

    The online round requires a prediction before the correct label is revealed.

    Fix: Keep the order as instance, prediction, correct label, and support for future predictions.

  • Calling an example only a test example or only a training example.

    In online learning, the same example first tests the current prediction and later contributes to future predictions after its label is obtained.

    Fix: Track the example's role over time: test before feedback, training after feedback.

  • Describing PAC learning and online learning as having the same schedule.

    PAC learning begins with a batch and then applies the learned hypothesis to new examples, whereas online learning processes examples through consecutive rounds.

    Fix: Compare the timing of training, prediction, and feedback rather than only comparing the names of the models.

Check the Sequence

MEDIUM

A learner receives one current instance, predicts its label, receives the correct label, and then uses the labeled instance when making a prediction about the next instance. Explain which event makes the current instance a test example, which event lets it serve as training information, and whether this schedule matches online learning or the PAC learning model.

Hints
  • Identify what the learner knows at the moment of prediction.
  • Look for the point at which the correct label becomes available.
  • PAC learning begins with a batch of training examples, while online learning repeats prediction and later feedback across rounds.

What do you think happens?

A current example has just been predicted, but its correct label has not yet been revealed. Is it already training information for future predictions?

  • Yes, because every current example is immediately training data
  • No, it first serves as a test example; it can support future predictions after its correct label is obtained
Reveal answer

Answer: No, it first serves as a test example; it can support future predictions after its correct label is obtained

The example tests the learner's current prediction before feedback. Once the correct label is available, the example can contribute to future predictions.

Round by Round

  1. Online learning is organized as a sequence of consecutive learning rounds.
  2. Each round presents an instance, requires a prediction, reveals the correct label, and uses that label to support future predictions.
  3. An example first tests the learner's current prediction and then becomes training information after its correct label is obtained.
  4. PAC learning separates a batch-based training stage from a later prediction stage, while online learning interweaves prediction and learning.

Key Takeaways

  • Online learning has no separate training phase followed by a separate prediction phase.
  • An online learning round proceeds from instance to prediction to correct label to support for future predictions.
  • The same example can move from a test role to a training role once its label is known.
  • PAC learning starts with a batch of training examples, whereas online learning processes examples through consecutive rounds.