Sequence Data
Timeseries data consists of measurements arranged in sequence at regular time intervals.
Why Order Matters
A sequence is more than a collection of values. It is a collection whose order carries information. Numerical timeseries data consists of measurements arranged in sequence at regular time intervals. A weather station, for example, records measurements, waits for a fixed interval, records them again, and continues this process over time. The position of each measurement in the sequence tells us when it occurred relative to the others.
When the order of observations helps describe the problem, the observations can be treated as sequence data rather than as unrelated records.
Weather Along a Time Axis
The weather dataset comes from a station at the Max Planck Institute for Biogeochemistry in Jena, Germany. It contains 14 weather-related quantities recorded every 10 minutes. The example uses data from 2009 through 2016, although the original data reaches back to 2003.
At each time point, the dataset can be viewed as a group of aligned measurements: the weather quantities recorded at that moment. These quantities include air temperature, atmospheric pressure, humidity, wind direction, and other weather measurements. The timestamp places that group on the time axis, while the measured quantities describe the weather at that point.
From History to Forecast
The temperature-forecasting task gives the model a few days of recent weather data and asks it to predict air temperature 24 hours in the future. This creates two different roles. The recent measurements are the historical input supplied as evidence. The later air-temperature measurement is the future target that the model must estimate.
Separating the input history from the forecast target makes the task precise. It is not an undefined question about weather. It is a sequence-to-future-value problem: use an ordered window of recent measurements to estimate a specified later value.
Labeling a Forecasting Window
Suppose a learner is shown a consecutive window of recent weather measurements and is asked for the air temperature 24 hours later. Which part is the input, and which part is the target?
Find the historical window: The few days of recent weather data form the evidence supplied to the model. They are the input sequence.
Find the later measurement: The air-temperature measurement at the specified future time is the value the model is asked to estimate.
Preserve the order: The measurements are used as an ordered history, not as an unordered collection of weather records.
Historical weather measurements are the input; the later air temperature 24 hours in the future is the target.
Records Become a Sequence
A single weather record contains measurements taken at one moment. The sequence view adds the relationship between that record and the records before and after it. Because the station records measurements repeatedly at a fixed interval, neighboring records are connected by their positions in time. A temperature value can therefore be interpreted as one point in an ordered history rather than only as an isolated number.
Treating weather records as isolated observations hides the ordered history that is used to frame the future prediction.
Language as a Sequence
A language model learns the statistical structure of language by modeling possible next tokens. A token can refer to a word or to a character. The model uses the preceding sequence to assign probabilities to possible next tokens.
This definition uses the same sequence idea in a different domain. In weather data, positions correspond to measurement times. In language, positions correspond to earlier and later tokens. The task is not necessarily to memorize one fixed continuation. The model represents statistical structure and uses the preceding sequence to determine which next tokens are more or less probable.
Generation Through Feedback
A trained language model can generate a sequence without writing the entire sequence in one step. It begins with a short piece of text called conditioning data. The model uses that initial sequence to predict a next token. The predicted token is appended to the sequence, and the enlarged sequence is supplied for another prediction.
This repeated feedback is the central mechanism. The first prediction has the conditioning data as context. The next prediction has the conditioning data plus the first generated token as context. Repeating the cycle can produce a sequence of arbitrary length.
Tracing Two Generation Steps
A generator starts with conditioning data and must produce a sequence one token at a time. What changes after each prediction?
Start: The model receives the short initial sequence, called conditioning data.
First prediction: The model produces a probability distribution for possible next tokens and uses it to produce one token.
Feedback: The generated token is attached to the conditioning data, creating an enlarged sequence.
Second prediction: The model predicts again, now using the enlarged sequence and therefore more context than in the first step.
Generation is a repeated cycle of predict, append, and predict again.
Character-Level Prediction
Language modeling is the general idea of predicting a next token from previous tokens. Character-level language modeling chooses individual characters as the prediction unit. In the LSTM character model described here, the model receives strings containing N characters and is trained to predict character N + 1.
The output of this LSTM character model is a softmax over all possible characters. That output forms a probability distribution for the next character. The model can then use the selected character as part of the enlarged sequence for the next prediction.
| Idea | Prediction unit | What happens next |
|---|---|---|
| Language modeling in general | A token, such as a word or character | The model assigns probabilities to possible next tokens. |
| Character-level language modeling | An individual character | The model predicts the next character from the preceding characters. |
Common Misunderstandings
Treating each weather record as independent.
The station records measurements at regular intervals, so the ordered history is part of the forecasting problem.
Fix:
Preserve the time order and view each record as one position in a weather sequence.Confusing the historical input with the forecast target.
The task uses recent weather measurements as input and asks for air temperature 24 hours in the future as the target.
Fix:
Label the recent window as historical input and the later air temperature as the future target.Assuming a language model generates an entire sentence in one step.
The described generation process predicts one next token, appends it, and predicts again.
Fix:
Trace the repeated feedback cycle: predict, append, and use the enlarged sequence as new context.Using character-level modeling and language modeling as if they were identical terms.
Language modeling concerns predicting the next token, while character-level modeling specifies that the token is an individual character.
Fix:
Treat character-level modeling as one particular choice of prediction unit.
Check Your Understanding
A weather station records several measurements every 10 minutes. A forecasting task gives a model a few days of recent data and asks it to estimate air temperature 24 hours later. Explain why this is sequence data, identify the input and target, and describe how the same sequence idea appears in language generation.
Hints
- Start by explaining what the regular 10-minute interval contributes.
- Separate the recent measurements from the later air-temperature value.
- For language generation, use the cycle predict, append, and predict again.
What do you think happens?
After a language model predicts one token and that token is appended to the sequence, what is supplied as context for the next prediction?
Reveal answer
Answer: The original conditioning data plus the newly generated token
The generated token is fed back into the sequence, so the next prediction uses the enlarged sequence and has more context than the previous prediction.
Key Takeaways
- Timeseries data consists of measurements arranged in order at regular time intervals.
- The weather dataset contains 14 weather-related quantities recorded every 10 minutes; the example uses data from 2009 through 2016.
- In the temperature task, recent weather measurements are the historical input and air temperature 24 hours later is the future target.
- A language model predicts possible next tokens from previous tokens.
- Sequence generation repeatedly predicts one token, appends it to the sequence, and feeds the enlarged sequence back into the model.
- Character-level language modeling is a specific form of language modeling in which each individual character is the prediction unit.
Key Takeaways
- Sequence data is ordered data whose positions represent regular measurement times or earlier and later tokens.
- The weather forecasting setup uses a recent historical window to predict air temperature 24 hours in the future.
- Weather records form a sequence because the station repeatedly records aligned measurements at fixed intervals.
- A language model predicts the next token from the preceding sequence.
- Generation uses feedback: predict one token, append it, and predict again.