Concepts / Sequence Data

Sequence Data

Timeseries data consists of measurements arranged in sequence at regular time intervals.

  • Programming

Why Order Matters

A sequence is more than a collection of values. It is a collection whose order carries information. Numerical timeseries data consists of measurements arranged in sequence at regular time intervals. A weather station, for example, records measurements, waits for a fixed interval, records them again, and continues this process over time. The position of each measurement in the sequence tells us when it occurred relative to the others.

When the order of observations helps describe the problem, the observations can be treated as sequence data rather than as unrelated records.

Weather Along a Time Axis

The weather dataset comes from a station at the Max Planck Institute for Biogeochemistry in Jena, Germany. It contains 14 weather-related quantities recorded every 10 minutes. The example uses data from 2009 through 2016, although the original data reaches back to 2003.

At each time point, the dataset can be viewed as a group of aligned measurements: the weather quantities recorded at that moment. These quantities include air temperature, atmospheric pressure, humidity, wind direction, and other weather measurements. The timestamp places that group on the time axis, while the measured quantities describe the weather at that point.

next intervalnext intervalnext intervalTime point 1timestamp + weathermeasurementsTime point 210 minutes laterTime point 310 minutes laterTime point 410 minutes later
How are weather measurements ordered along a regular time axis, and what does each position represent?

From History to Forecast

The temperature-forecasting task gives the model a few days of recent weather data and asks it to predict air temperature 24 hours in the future. This creates two different roles. The recent measurements are the historical input supplied as evidence. The later air-temperature measurement is the future target that the model must estimate.

Separating the input history from the forecast target makes the task precise. It is not an undefined question about weather. It is a sequence-to-future-value problem: use an ordered window of recent measurements to estimate a specified later value.

earlier to laterearlier to laterforecastRecent measurement 1historical inputRecent measurement 2historical inputRecent measurement 3historical inputAir temperature24 hours in the future
Which past temperature measurements are used as input, and which later measurement is the future target?

Labeling a Forecasting Window

Suppose a learner is shown a consecutive window of recent weather measurements and is asked for the air temperature 24 hours later. Which part is the input, and which part is the target?

Find the historical window: The few days of recent weather data form the evidence supplied to the model. They are the input sequence.

Find the later measurement: The air-temperature measurement at the specified future time is the value the model is asked to estimate.

Preserve the order: The measurements are used as an ordered history, not as an unordered collection of weather records.

Historical weather measurements are the input; the later air temperature 24 hours in the future is the target.

Records Become a Sequence

A single weather record contains measurements taken at one moment. The sequence view adds the relationship between that record and the records before and after it. Because the station records measurements repeatedly at a fixed interval, neighboring records are connected by their positions in time. A temperature value can therefore be interpreted as one point in an ordered history rather than only as an isolated number.

time passestime passestime passesWeather record 1one measurement timeWeather record 2next measurement timeWeather record 3next measurement timeWeather record 4next measurement time
How are neighboring weather records connected through time, and what is lost when each record is treated as independent?

Treating weather records as isolated observations hides the ordered history that is used to frame the future prediction.

Language as a Sequence

A language model learns the statistical structure of language by modeling possible next tokens. A token can refer to a word or to a character. The model uses the preceding sequence to assign probabilities to possible next tokens.

This definition uses the same sequence idea in a different domain. In weather data, positions correspond to measurement times. In language, positions correspond to earlier and later tokens. The task is not necessarily to memorize one fixed continuation. The model represents statistical structure and uses the preceding sequence to determine which next tokens are more or less probable.

Generation Through Feedback

A trained language model can generate a sequence without writing the entire sequence in one step. It begins with a short piece of text called conditioning data. The model uses that initial sequence to predict a next token. The predicted token is appended to the sequence, and the enlarged sequence is supplied for another prediction.

This repeated feedback is the central mechanism. The first prediction has the conditioning data as context. The next prediction has the conditioning data plus the first generated token as context. Repeating the cycle can produce a sequence of arbitrary length.

provide contextpredictappendfeed backpredict againConditioning textinitial sequenceNext-token predictionprobabilitiesToken 1generated outputExpanded sequenceconditioning text + Token 1Next-token predictionmore contextToken 2generated output
How does a language model use previous tokens, generate one new token, feed it back into the input, and continue producing the sequence?

Tracing Two Generation Steps

A generator starts with conditioning data and must produce a sequence one token at a time. What changes after each prediction?

Start: The model receives the short initial sequence, called conditioning data.

First prediction: The model produces a probability distribution for possible next tokens and uses it to produce one token.

Feedback: The generated token is attached to the conditioning data, creating an enlarged sequence.

Second prediction: The model predicts again, now using the enlarged sequence and therefore more context than in the first step.

Generation is a repeated cycle of predict, append, and predict again.

Character-Level Prediction

Language modeling is the general idea of predicting a next token from previous tokens. Character-level language modeling chooses individual characters as the prediction unit. In the LSTM character model described here, the model receives strings containing N characters and is trained to predict character N + 1.

The output of this LSTM character model is a softmax over all possible characters. That output forms a probability distribution for the next character. The model can then use the selected character as part of the enlarged sequence for the next prediction.

IdeaPrediction unitWhat happens next
Language modeling in generalA token, such as a word or characterThe model assigns probabilities to possible next tokens.
Character-level language modelingAn individual characterThe model predicts the next character from the preceding characters.

Common Misunderstandings

  • Treating each weather record as independent.

    The station records measurements at regular intervals, so the ordered history is part of the forecasting problem.

    Fix: Preserve the time order and view each record as one position in a weather sequence.

  • Confusing the historical input with the forecast target.

    The task uses recent weather measurements as input and asks for air temperature 24 hours in the future as the target.

    Fix: Label the recent window as historical input and the later air temperature as the future target.

  • Assuming a language model generates an entire sentence in one step.

    The described generation process predicts one next token, appends it, and predicts again.

    Fix: Trace the repeated feedback cycle: predict, append, and use the enlarged sequence as new context.

  • Using character-level modeling and language modeling as if they were identical terms.

    Language modeling concerns predicting the next token, while character-level modeling specifies that the token is an individual character.

    Fix: Treat character-level modeling as one particular choice of prediction unit.

Check Your Understanding

MEDIUM

A weather station records several measurements every 10 minutes. A forecasting task gives a model a few days of recent data and asks it to estimate air temperature 24 hours later. Explain why this is sequence data, identify the input and target, and describe how the same sequence idea appears in language generation.

Hints
  • Start by explaining what the regular 10-minute interval contributes.
  • Separate the recent measurements from the later air-temperature value.
  • For language generation, use the cycle predict, append, and predict again.

What do you think happens?

After a language model predicts one token and that token is appended to the sequence, what is supplied as context for the next prediction?

  • Only the newly generated token
  • The original conditioning data plus the newly generated token
  • Only the original conditioning data
  • The future target from the weather dataset
Reveal answer

Answer: The original conditioning data plus the newly generated token

The generated token is fed back into the sequence, so the next prediction uses the enlarged sequence and has more context than the previous prediction.

Key Takeaways

  1. Timeseries data consists of measurements arranged in order at regular time intervals.
  2. The weather dataset contains 14 weather-related quantities recorded every 10 minutes; the example uses data from 2009 through 2016.
  3. In the temperature task, recent weather measurements are the historical input and air temperature 24 hours later is the future target.
  4. A language model predicts possible next tokens from previous tokens.
  5. Sequence generation repeatedly predicts one token, appends it to the sequence, and feeds the enlarged sequence back into the model.
  6. Character-level language modeling is a specific form of language modeling in which each individual character is the prediction unit.

Key Takeaways

  • Sequence data is ordered data whose positions represent regular measurement times or earlier and later tokens.
  • The weather forecasting setup uses a recent historical window to predict air temperature 24 hours in the future.
  • Weather records form a sequence because the station repeatedly records aligned measurements at fixed intervals.
  • A language model predicts the next token from the preceding sequence.
  • Generation uses feedback: predict one token, append it, and predict again.