Recurrent Neural Networks
A language model learns the statistical structure of language by modeling possible next tokens.
From Prediction to Generation
A language model learns the statistical structure of language by modeling possible next tokens. A token can be a word or a character. The model does not need to produce an entire sentence in one step. Instead, it predicts what should come next, adds that prediction to the sequence, and predicts again. Repeating this cycle turns a next-token predictor into a sequence generator.
The central mechanism is feedback: each predicted token becomes part of the sequence used for the next prediction.
Tracing the Generation Loop
Growing a Sequence
Suppose a trained language model receives the conditioning data "The". Describe the conceptual steps by which it can generate more text.
Start with context: The short text "The" is the conditioning data. It gives the model the initial sequence from which to make a prediction.
Predict one token: The model uses the conditioning data to assign probabilities to possible next tokens and selects one next token.
Append the result: The selected token is attached to the existing sequence, producing an enlarged sequence.
Predict again: The model uses the enlarged sequence for another next-token prediction. The cycle can continue repeatedly.
Generation proceeds incrementally: condition, predict, append, and predict again. The example does not require the model to write the whole sentence in one step.
Why Sequence Order Matters
Sequence data is challenging because the meaning assigned to a new item can be influenced by the items that came before it. When reading a sentence one word at a time, earlier words can affect how a later word is understood. A sequence is therefore not merely a collection of separate inputs: position and previous information can affect the interpretation of the current item.
A recurrent neural network addresses this challenge by processing information incrementally while maintaining an internal model. When another item arrives, the network updates that internal model, allowing information from earlier items to remain available for processing what comes next.
Carrying State Forward
The defining procedural change is that information arrives incrementally and an internal model is maintained between steps. The network does not treat every item as an isolated input. Instead, processing one item updates the internal model, and that updated model is available when the next item arrives.
A recurrent network's state is an internal model that is updated as sequence items arrive. This carried context lets later items be processed using information from earlier items.
Feedforward and Recurrent Processing
| Processing approach | Treatment of inputs | Use of earlier information |
|---|---|---|
| Feedforward | Processes inputs independently | Does not maintain state between inputs |
| Recurrent | Processes information incrementally | Maintains and updates an internal model carrying earlier information |
Character-Level Modeling
Language modeling is a general idea: predict a possible next token from the preceding sequence. Character-level modeling chooses individual characters as those tokens. In the LSTM character-model setup described by the source, the model receives a string containing N characters and is trained to predict character N + 1. Its output is a probability distribution over possible next characters.
| Aspect | Language modeling in general | Character-level modeling |
|---|---|---|
| Prediction unit | A token, which may be a word or a character | An individual character |
| Prediction target | A possible next token | Character N + 1 after a string of N characters |
| Model output | Probabilities for possible next tokens | A probability distribution over possible next characters |
When explaining a language model, identify the token unit first. Saying that a model predicts the next token is general; saying that it predicts the next character describes a character-level model.
Mistakes About Recurrent Networks
Assuming that a language model generates a whole sentence in one step.
The described generation process predicts one next token, appends it, and predicts again.
Fix:
Think in repeated cycles: conditioning data, next-token prediction, append, and another prediction.Treating character-level modeling as the definition of language modeling.
A token can be a word or a character. Character-level modeling is one choice of token unit.
Fix:
First define language modeling as next-token prediction, then specify whether the token is a character or another unit.Assuming that presenting inputs in sequence automatically gives a network memory.
Feedforward networks process inputs independently because they do not maintain state between inputs.
Fix:
Look for an internal model that is maintained and updated as each item arrives.Ignoring earlier items when interpreting a later sequence item.
Previous information can influence how a new word is understood.
Fix:
Ask what context has been carried forward before processing the current item.
Practice the Mechanism
Explain how a trained language model could generate a sequence beginning with a short conditioning text. In your explanation, use the terms conditioning data, next-token prediction, append, feedback, and internal model. Then state one difference between feedforward and recurrent processing.
Hints
- Begin with the initial text rather than with a complete sentence.
- Describe what happens to the predicted token before the next prediction.
- For the processing comparison, focus on whether state is maintained between inputs.
What do you think happens?
After a model predicts one token from its conditioning data, what happens before the next prediction?
Reveal answer
Answer: The predicted token is appended to the sequence.
The enlarged sequence then provides more context for the next prediction, and the cycle can continue.
What to Remember
- A language model learns statistical structure by predicting possible next tokens from preceding tokens.
- Generation uses feedback: predict one token, append it to the sequence, and predict again.
- Character-level modeling uses individual characters as tokens and predicts the next character.
- Feedforward networks process inputs independently, while recurrent processing maintains and updates an internal model.
- Earlier sequence items can influence how later items are interpreted or predicted.
Key Takeaways
- A language model predicts possible next tokens from previous tokens.
- A sequence generator repeatedly predicts, appends the result, and predicts again.
- Recurrent processing carries an updated internal model from earlier sequence items to later ones.
- Feedforward processing treats inputs independently because it does not maintain state between inputs.
- Character-level modeling is a specific form of language modeling in which each character is a token.