Sequence Modeling
A language model learns the statistical structure of language by modeling possible next tokens.
From Context to Continuation
A language model learns the statistical structure of language by modeling possible next tokens. It receives a sequence of previous tokens and uses that sequence to assign probabilities to possible next tokens. The word token is general: depending on the model, a token can be a word or a character.
A language model predicts what token could come next given the tokens that came before it.
A First Generation Trace
Generation does not require the model to write an entire sentence in one step. Instead, generation begins with a short piece of text. That starting text is the conditioning data. The model predicts a next token, the result is appended to the sequence, and the enlarged sequence becomes the context for another prediction.
Growing a Sequence
Trace the feedback cycle for a character-level generator that begins with the conditioning data "ca" and predicts the next character three times.
Starting context: The generator begins with the conditioning characters "ca". These characters are the initial sequence supplied to the trained model.
First prediction: The model uses "ca" to produce a probability distribution over possible next characters. Suppose the selected output is "t".
Append the result: The selected character is attached to the sequence, producing "cat".
Second prediction: The model now uses the enlarged sequence "cat" as context. Suppose the selected output is a space character.
Append again: The sequence becomes "cat ". The next prediction can use more context than the first prediction did.
Third prediction: The model uses "cat " as its current context, produces another probability distribution, and selects the next character.
The sequence is generated through repeated prediction and appending. The model does not need to produce the whole sequence in one step.
What do you think happens?
After a model predicts and appends one token, what sequence does it use for the next prediction?
Reveal answer
Answer: The model uses the enlarged sequence: the conditioning data together with the newly appended token.
The central generation cycle is predict, append the result, and predict again. Appending gives the next prediction more preceding context.
Feedback as the Central Mechanism
The defining mechanism of sequence generation is feedback. The initial text is called conditioning data. The trained model uses that data to produce a next-character or next-word prediction. The output is then added to the sequence. Because the sequence is now longer, the next prediction has more context than the previous prediction. This cycle can continue many times and can produce sequences of arbitrary length.
Character-Level Modeling
Character-level language modeling is one particular form of language modeling. It chooses the character as the prediction unit. In the LSTM character model described in the source material, the model receives strings containing N characters and is trained to predict character N + 1. Its output is a softmax over all possible characters, forming a probability distribution for the next character.
| Aspect | Language modeling in general | Character-level language modeling |
|---|---|---|
| Prediction unit | A token, where a token can refer to a word or a character | An individual character |
| Input context | Previous tokens | A string containing N characters |
| Prediction | A distribution over possible next tokens | A distribution over possible next characters |
| Generation step | Predict and append one token | Predict and append one character |
The distinction is about the prediction unit. Character-level modeling does not replace the general idea of sequence modeling; it applies that idea with individual characters as the tokens.
Mistakes About Generation
Treating a language model as a system that must produce an entire sentence in one step.
The described generation mechanism predicts a next token, appends it, and predicts again.
Fix:
Trace generation as a repeated cycle in which the sequence grows one token at a time.Ignoring the role of conditioning data.
The initial text supplies the context used by the trained model for its first prediction.
Fix:
Identify the initial text as conditioning data before tracing the first prediction.Assuming that generated output is not used again.
The newly predicted token is added to the sequence, so later predictions use the enlarged context.
Fix:
After every prediction, append the result and feed the enlarged sequence into the next step.Equating language modeling only with word prediction.
A token can refer to a word or a character, and character-level modeling is a form of language modeling.
Fix:
Ask what unit the model treats as a token before describing its prediction task.
When explaining or analyzing a sequence generator, write down three things at each step: the current sequence, the predicted next token, and the enlarged sequence after appending. This makes the feedback mechanism explicit.
Check Your Understanding
A character-level generator starts with the conditioning data "go". It predicts "!" as the next character. Describe the exact context that should be used for the following prediction, and explain why.
Hints
- First write the sequence after the predicted character is appended.
- Then identify which sequence is available as context for the next prediction.
Explain the difference between these two statements: "The model predicts the next token" and "The model generates the whole sequence." Your answer should describe the repeated feedback process that connects them.
Hints
- The first statement describes one prediction step.
- The second statement describes what happens when that step is repeated after each appended output.
Sequence Modeling in One View
- A language model represents the statistical structure of language by modeling possible next tokens from previous tokens.
- Generation begins with conditioning data and proceeds by predicting a token, appending it, and predicting again.
- Feedback is central: every newly generated token becomes part of the context for a later prediction.
- A token may be a word or a character, so character-level language modeling is a specific form of language modeling.
- An LSTM character model predicts character N + 1 from a string containing N characters and produces a probability distribution over possible next characters.
Key Takeaways
- A language model predicts possible next tokens from the preceding sequence.
- A trained model generates longer sequences through repeated prediction and appending.
- Conditioning data provides the starting context, while feedback supplies increasingly enlarged context.
- Character-level language modeling uses individual characters as its tokens and predicts one next character at a time.