Concepts / Recurrent Neural Networks

Recurrent Neural Networks

A language model learns the statistical structure of language by modeling possible next tokens.

  • Programming

From Prediction to Generation

A language model learns the statistical structure of language by modeling possible next tokens. A token can be a word or a character. The model does not need to produce an entire sentence in one step. Instead, it predicts what should come next, adds that prediction to the sequence, and predicts again. Repeating this cycle turns a next-token predictor into a sequence generator.

conditionappendpredict againappend and repeatconditioning datainitial textnext tokenprediction 1enlarged sequencecontext plus tokennext tokenprediction 2
How does a language model use previous tokens and its predicted next token to generate an entire sequence step by step?

The central mechanism is feedback: each predicted token becomes part of the sequence used for the next prediction.

Tracing the Generation Loop

Growing a Sequence

Suppose a trained language model receives the conditioning data "The". Describe the conceptual steps by which it can generate more text.

Start with context: The short text "The" is the conditioning data. It gives the model the initial sequence from which to make a prediction.

Predict one token: The model uses the conditioning data to assign probabilities to possible next tokens and selects one next token.

Append the result: The selected token is attached to the existing sequence, producing an enlarged sequence.

Predict again: The model uses the enlarged sequence for another next-token prediction. The cycle can continue repeatedly.

Generation proceeds incrementally: condition, predict, append, and predict again. The example does not require the model to write the whole sentence in one step.

conditionspredictsis appended tofeeds the next predictionconditioning datainitial sequencelanguage modelnext-token probabilitiespredicted tokennew outputexpanded sequencecontext plus output
How does the initial context condition the first prediction, and how does each generated token become input for the next prediction?

Why Sequence Order Matters

Sequence data is challenging because the meaning assigned to a new item can be influenced by the items that came before it. When reading a sentence one word at a time, earlier words can affect how a later word is understood. A sequence is therefore not merely a collection of separate inputs: position and previous information can affect the interpretation of the current item.

comes beforeinfluencesis interpreted with contextearlier wordprevious informationcurrent wordnew itemcurrentinterpretationuses prior context
How does information from an earlier token influence the network's interpretation or prediction at a later position?

A recurrent neural network addresses this challenge by processing information incrementally while maintaining an internal model. When another item arrives, the network updates that internal model, allowing information from earlier items to remain available for processing what comes next.

Carrying State Forward

The defining procedural change is that information arrives incrementally and an internal model is maintained between steps. The network does not treat every item as an isolated input. Instead, processing one item updates the internal model, and that updated model is available when the next item arrives.

update with itemis processedcarry forward and updateis processedinternal modelbefore new itemsequence itemarrivesupdated internalmodelearlier information carriedforwardnext sequence itemarrives laterupdated internalmodelcontext updated again
What changes inside the network as each new token is processed, and how is information from earlier tokens carried forward?

A recurrent network's state is an internal model that is updated as sequence items arrive. This carried context lets later items be processed using information from earlier items.

Feedforward and Recurrent Processing

Processing approachTreatment of inputsUse of earlier information
FeedforwardProcesses inputs independentlyDoes not maintain state between inputs
RecurrentProcesses information incrementallyMaintains and updates an internal model carrying earlier information
updatesinformsinput 1independentinput 1updates stateinput 2independentinternal modelcarried forwardinput 2uses state
What is the difference between processing each input independently and processing inputs while carrying information from one step to the next?

Character-Level Modeling

Language modeling is a general idea: predict a possible next token from the preceding sequence. Character-level modeling chooses individual characters as those tokens. In the LSTM character-model setup described by the source, the model receives a string containing N characters and is trained to predict character N + 1. Its output is a probability distribution over possible next characters.

AspectLanguage modeling in generalCharacter-level modeling
Prediction unitA token, which may be a word or a characterAn individual character
Prediction targetA possible next tokenCharacter N + 1 after a string of N characters
Model outputProbabilities for possible next tokensA probability distribution over possible next characters

When explaining a language model, identify the token unit first. Saying that a model predicts the next token is general; saying that it predicts the next character describes a character-level model.

Mistakes About Recurrent Networks

  • Assuming that a language model generates a whole sentence in one step.

    The described generation process predicts one next token, appends it, and predicts again.

    Fix: Think in repeated cycles: conditioning data, next-token prediction, append, and another prediction.

  • Treating character-level modeling as the definition of language modeling.

    A token can be a word or a character. Character-level modeling is one choice of token unit.

    Fix: First define language modeling as next-token prediction, then specify whether the token is a character or another unit.

  • Assuming that presenting inputs in sequence automatically gives a network memory.

    Feedforward networks process inputs independently because they do not maintain state between inputs.

    Fix: Look for an internal model that is maintained and updated as each item arrives.

  • Ignoring earlier items when interpreting a later sequence item.

    Previous information can influence how a new word is understood.

    Fix: Ask what context has been carried forward before processing the current item.

Practice the Mechanism

MEDIUM

Explain how a trained language model could generate a sequence beginning with a short conditioning text. In your explanation, use the terms conditioning data, next-token prediction, append, feedback, and internal model. Then state one difference between feedforward and recurrent processing.

Hints
  • Begin with the initial text rather than with a complete sentence.
  • Describe what happens to the predicted token before the next prediction.
  • For the processing comparison, focus on whether state is maintained between inputs.

What do you think happens?

After a model predicts one token from its conditioning data, what happens before the next prediction?

  • The predicted token is appended to the sequence.
  • The model discards the conditioning data.
  • The model produces the entire remaining sentence at once.
Reveal answer

Answer: The predicted token is appended to the sequence.

The enlarged sequence then provides more context for the next prediction, and the cycle can continue.

What to Remember

  1. A language model learns statistical structure by predicting possible next tokens from preceding tokens.
  2. Generation uses feedback: predict one token, append it to the sequence, and predict again.
  3. Character-level modeling uses individual characters as tokens and predicts the next character.
  4. Feedforward networks process inputs independently, while recurrent processing maintains and updates an internal model.
  5. Earlier sequence items can influence how later items are interpreted or predicted.

Key Takeaways

  • A language model predicts possible next tokens from previous tokens.
  • A sequence generator repeatedly predicts, appends the result, and predicts again.
  • Recurrent processing carries an updated internal model from earlier sequence items to later ones.
  • Feedforward processing treats inputs independently because it does not maintain state between inputs.
  • Character-level modeling is a specific form of language modeling in which each character is a token.