Sampling from Probability Distributions
Character-level generation learns the relationship between a fixed-length character context and the character that follows it.
From Text to Training Pairs
Character-level generation begins with a narrow question: given the characters seen so far, which character should come next? The model does not generate an entire word or sentence in one operation. Training therefore converts a continuous text into many fixed-length contexts, with each context paired with the single character that follows it.
One Context and Its Target
Suppose a training example uses a context length of three characters and begins with the text abcdef.
Select the context: The first fixed-length context is abc.
Find the following character: The character immediately after abc is d, so d is the target for this example.
Shift the window: Moving the window forward by one character produces bcd, whose following character is e.
Repeat: The next window is cde, paired with f. The contexts overlap because each new window reuses most of the previous characters.
The training pairs are abc to d, bcd to e, and cde to f.
In the source implementation, the input array x is three-dimensional. Its dimensions represent the number of sequences, the sequence length maxlen, and the number of unique characters. The target array y contains the one-hot-encoded character that follows each input sequence. This alignment is the central data contract: x supplies the context, and y identifies the character the model should learn to predict.
The Prediction Pipeline
After the training pairs have been prepared, the model learns the relationship between each fixed-length context and its following character. The model combines an LSTM with a Dense softmax classifier. The LSTM processes the character sequence, while the Dense softmax classifier produces probabilities for possible next characters. The output size matches the character vocabulary, so the prediction represents the available character choices.
A character vocabulary is the set of unique characters represented by the model. Because the Dense softmax classifier has one output position for each available character, its output size must match that vocabulary.
One Generation Step
Once training is complete, generation starts with a seed snippet. The model predicts a distribution for the next character. A sampling procedure chooses one character from that distribution. The chosen character is appended to the generated text, and the oldest character is removed from the fixed-length context. The resulting window becomes the input for the next prediction.
Tracing Several Generation Steps
Assume the current context is abc and the model samples d as the next character.
Start: The model receives the current fixed-length context abc.
Predict: The model produces a probability distribution over the characters in its vocabulary.
Sample: One character is selected from that distribution; in this example, the selected character is d.
Append: The generated text receives d after the existing seed or passage.
Shift: The oldest character is removed from the fixed-length context, leaving bcd as the next context.
The next prediction uses bcd, so the sampled result becomes part of the context that drives later predictions.
This process is autoregressive: each sampled character becomes part of the context for the next prediction. An early sampled choice can therefore influence the contexts used in later steps.
Temperature and Variety
Temperature changes the probability distribution used for sampling. It does not create a separate trained model. Instead, different temperature settings provide alternative ways to sample from the same model predictions, allowing you to compare more concentrated output with more varied output.
| Setting | Role | What to compare |
|---|---|---|
| 0.2 | Temperature used for sampling | How concentrated the resulting output is |
| 0.5 | Temperature used for sampling | How the output differs from other settings |
| 1.0 | Temperature used for sampling | How varied the resulting output is |
| 1.2 | Temperature used for sampling | How the sampling choice changes the passage |
The source loop compares these settings while using the same trained model predictions.
Debugging the Generation Pipeline
A generated passage can look wrong for more than one reason. Separate the pipeline into data preparation, model construction, and sampling. This classification prevents a sampling symptom from being mistaken for a training-data problem, or a vocabulary mismatch from being blamed on temperature.
Pairing a context with the wrong following character.
The target array is supposed to identify the character that follows each input sequence. Misalignment breaks the central data contract.
Fix:
Check that every input sequence and its target come from the same position in the original text.Using an output layer that does not represent the full character vocabulary.
The model cannot represent every possible character in its vocabulary.
Fix:
Verify that the Dense softmax output size matches the number of unique characters.Changing the temperature when the real problem is the sampling conversion.
Temperature changes the sampling distribution, but it does not repair an incorrect probability-to-index-to-character conversion.
Fix:
Inspect the complete conversion from predicted probabilities to a sampled index and then back to a character.Forgetting that generation changes its own context.
Generation is autoregressive. Each sampled character should become part of the context for the next prediction.
Fix:
Append the sampled character and remove the oldest character from the fixed-length context before predicting again.
- First check whether each input sequence is paired with the character that actually follows it.
- Then check whether the Dense softmax output represents the full character vocabulary.
- Finally check the conversion from predicted probabilities to a sampled index and then back to a character.
Practice the Trace
A model uses a context length of three characters. The current context is xyz, and the sampled next character is q. Describe the generated-text update and identify the context used for the next prediction.
Hints
- The sampled character is appended to the generated text.
- The next context must remain three characters long.
- Remove the oldest character after appending the sampled character.
A generated passage looks unrelated to the training corpus. Organize your investigation into three checks: one for data preparation, one for model construction, and one for sampling.
Hints
- For data preparation, inspect the relationship between each context and its following target.
- For model construction, inspect whether the output covers the full character vocabulary.
- For sampling, inspect the conversion from probabilities to an index and then to a character.
What do you think happens?
After the context abc produces the sampled character d, what fixed-length context is used next if the context length is three?
Reveal answer
Answer: bcd
The sampled character is appended, and the oldest character is removed from the fixed-length context. Starting from abc and adding d leaves bcd.
Key Takeaways
- Training converts continuous text into overlapping fixed-length character contexts and the characters that follow them.
- The input array supplies context sequences, while the one-hot-encoded target identifies the next character.
- An LSTM processes the sequence, and a vocabulary-sized Dense softmax classifier produces next-character probabilities.
- Generation repeatedly predicts, samples, appends the chosen character, and shifts the context window.
- Temperature changes how the same model predictions are sampled, while debugging should separately examine data preparation, model construction, and sampling.
Key Takeaways
- Character-level generation learns which character follows a fixed-length context.
- Overlapping contexts and correctly aligned next-character targets form the training data contract.
- The LSTM and Dense softmax layers turn a context into a probability distribution over the character vocabulary.
- Autoregressive generation feeds each sampled character into the next context.
- Temperature changes sampling behavior, and generation failures can be classified as data, model, or sampling problems.