Concepts / Sequence Modeling

Sequence Modeling

A language model learns the statistical structure of language by modeling possible next tokens.

  • Programming

From Context to Continuation

A language model learns the statistical structure of language by modeling possible next tokens. It receives a sequence of previous tokens and uses that sequence to assign probabilities to possible next tokens. The word token is general: depending on the model, a token can be a word or a character.

A language model predicts what token could come next given the tokens that came before it.

provideassignPrevious tokenscontextLanguage modellearned structureNext-tokenprobabilitiespossible continuations
How does a language model map previous tokens to a prediction of the next token?

A First Generation Trace

Generation does not require the model to write an entire sentence in one step. Instead, generation begins with a short piece of text. That starting text is the conditioning data. The model predicts a next token, the result is appended to the sequence, and the enlarged sequence becomes the context for another prediction.

Growing a Sequence

Trace the feedback cycle for a character-level generator that begins with the conditioning data "ca" and predicts the next character three times.

Starting context: The generator begins with the conditioning characters "ca". These characters are the initial sequence supplied to the trained model.

First prediction: The model uses "ca" to produce a probability distribution over possible next characters. Suppose the selected output is "t".

Append the result: The selected character is attached to the sequence, producing "cat".

Second prediction: The model now uses the enlarged sequence "cat" as context. Suppose the selected output is a space character.

Append again: The sequence becomes "cat ". The next prediction can use more context than the first prediction did.

Third prediction: The model uses "cat " as its current context, produces another probability distribution, and selects the next character.

The sequence is generated through repeated prediction and appending. The model does not need to produce the whole sequence in one step.

predict and appendpredict and appendpredict and appendcainitial sequencecatappend tcatappend spacenext sequenceappend another prediction
How does each predicted token become part of the context for the next prediction?

What do you think happens?

After a model predicts and appends one token, what sequence does it use for the next prediction?

  • Only the newly predicted token
  • The original conditioning data and the newly appended token
  • A completely unrelated sequence
Reveal answer

Answer: The model uses the enlarged sequence: the conditioning data together with the newly appended token.

The central generation cycle is predict, append the result, and predict again. Appending gives the next prediction more preceding context.

Feedback as the Central Mechanism

The defining mechanism of sequence generation is feedback. The initial text is called conditioning data. The trained model uses that data to produce a next-character or next-word prediction. The output is then added to the sequence. Because the sequence is now longer, the next prediction has more context than the previous prediction. This cycle can continue many times and can produce sequences of arbitrary length.

conditionproduceappendfeed back as contextConditioning datainitial tokensTrained modelnext-token probabilitiesPredicted tokennew outputEnlarged sequencemore context
How do the initial conditioning tokens and generated tokens flow back into the model?
modelbasis for selectingattachuse as next contextPrevious sequenceconditioning or enlargedcontextPossible next tokensprobability distributionNext outputselected tokenAppended sequencenew context
How does modeling possible next tokens become an ongoing sequence-generation process?

Character-Level Modeling

Character-level language modeling is one particular form of language modeling. It chooses the character as the prediction unit. In the LSTM character model described in the source material, the model receives strings containing N characters and is trained to predict character N + 1. Its output is a softmax over all possible characters, forming a probability distribution for the next character.

AspectLanguage modeling in generalCharacter-level language modeling
Prediction unitA token, where a token can refer to a word or a characterAn individual character
Input contextPrevious tokensA string containing N characters
PredictionA distribution over possible next tokensA distribution over possible next characters
Generation stepPredict and append one tokenPredict and append one character
predictpredictPrevious tokensword or character unitsN characterscharacter stringNext tokenpossible-token distributionNext characterpossible-characterdistribution
What changes when the model predicts characters instead of tokens more generally?

The distinction is about the prediction unit. Character-level modeling does not replace the general idea of sequence modeling; it applies that idea with individual characters as the tokens.

Mistakes About Generation

  • Treating a language model as a system that must produce an entire sentence in one step.

    The described generation mechanism predicts a next token, appends it, and predicts again.

    Fix: Trace generation as a repeated cycle in which the sequence grows one token at a time.

  • Ignoring the role of conditioning data.

    The initial text supplies the context used by the trained model for its first prediction.

    Fix: Identify the initial text as conditioning data before tracing the first prediction.

  • Assuming that generated output is not used again.

    The newly predicted token is added to the sequence, so later predictions use the enlarged context.

    Fix: After every prediction, append the result and feed the enlarged sequence into the next step.

  • Equating language modeling only with word prediction.

    A token can refer to a word or a character, and character-level modeling is a form of language modeling.

    Fix: Ask what unit the model treats as a token before describing its prediction task.

When explaining or analyzing a sequence generator, write down three things at each step: the current sequence, the predicted next token, and the enlarged sequence after appending. This makes the feedback mechanism explicit.

Check Your Understanding

EASY

A character-level generator starts with the conditioning data "go". It predicts "!" as the next character. Describe the exact context that should be used for the following prediction, and explain why.

Hints
  • First write the sequence after the predicted character is appended.
  • Then identify which sequence is available as context for the next prediction.
MEDIUM

Explain the difference between these two statements: "The model predicts the next token" and "The model generates the whole sequence." Your answer should describe the repeated feedback process that connects them.

Hints
  • The first statement describes one prediction step.
  • The second statement describes what happens when that step is repeated after each appended output.

Sequence Modeling in One View

  1. A language model represents the statistical structure of language by modeling possible next tokens from previous tokens.
  2. Generation begins with conditioning data and proceeds by predicting a token, appending it, and predicting again.
  3. Feedback is central: every newly generated token becomes part of the context for a later prediction.
  4. A token may be a word or a character, so character-level language modeling is a specific form of language modeling.
  5. An LSTM character model predicts character N + 1 from a string containing N characters and produces a probability distribution over possible next characters.

Key Takeaways

  • A language model predicts possible next tokens from the preceding sequence.
  • A trained model generates longer sequences through repeated prediction and appending.
  • Conditioning data provides the starting context, while feedback supplies increasingly enlarged context.
  • Character-level language modeling uses individual characters as its tokens and predicts one next character at a time.