Token Embeddings
Sequence data includes text, timeseries, and other ordered inputs.
From Sequence to Tensor
A sequence is ordered data. Text, timeseries, and other ordered inputs all have meaningful positions. A neural network cannot receive raw text directly; it needs a numeric tensor. Sequence processing therefore begins before the model runs: the original sequence must first be converted into a numerical representation.
The key transformation is not simply “text becomes numbers.” The ordered units must be identified first, then each unit must receive a numeric representation, and those representations must remain arranged as a sequence.
Tokenization and Vectorization
Tokenization chooses the units in a sequence. For text, the units can be words or characters. Tokenization does not yet provide the neural network with the vectors it needs; it identifies the pieces that will later be represented numerically.
Vectorization turns the generated tokens into numeric representations. The source describes two broad ways to associate vectors with tokens: one-hot encoding and token embedding. Token embedding is typically used for words and is also called word embedding. After vectorization, the vectors are packed into sequence tensors that can be given to a deep neural network.
Following One Short Text Sequence
Trace the conceptual conversion of the ordered text “birds fly” into an input suitable for a neural network.
Choose units: Treat the two words as the tokens “birds” and “fly.” This is tokenization: deciding which units make up the sequence.
Assign representations: Give “birds” and “fly” numeric vectors through a vectorization method such as token embedding. The exact vector values are not important for this conceptual trace.
Preserve order: Keep the vector for “birds” before the vector for “fly,” because the input is an ordered sequence.
Pack the sequence: Place the ordered vectors into a sequence tensor, which can then be fed into a deep neural network.
The text has become an ordered sequence of numeric vectors rather than remaining raw text.
Embedding Lookup
A token embedding associates a token with a numeric vector. Conceptually, you can think of a lookup from each token to its vector. Applying that lookup across an ordered token sequence produces an ordered sequence of vectors. The resulting collection is the representation that can be packed into a sequence tensor.
Sliding Through a Sequence
Once a sequence has been represented as a numeric tensor, a model designed for sequences can process it. A 1D convnet applies the one-dimensional convolutional approach to sequence processing. Its defining operation is a one-dimensional convolution applied across the ordered sequence.
The important conceptual connection is this: token embeddings provide numeric vectors at ordered positions, and a 1D convnet applies a sequence-processing operation to those positions. The model therefore works on the numerical sequence representation rather than on raw words or characters.
Convnets and Recurrent Networks
The source identifies two fundamental deep-learning approaches for sequence processing: 1D convnets and recurrent neural networks. A 1D convnet uses the one-dimensional convolutional approach on the sequence. A recurrent neural network is the other broad approach named for sequence processing. At this level, the main distinction is the model family being applied after the sequence has been converted into numeric tensors.
When choosing terminology, do not call every step “embedding.” Embedding describes the numeric representation associated with tokens. A 1D convnet or recurrent neural network describes the sequence-processing model that can operate on the resulting tensor.
Where Sequence Models Fit
Sequence-processing models are appropriate when order is part of the input. Text is one example: it can be treated as a sequence of words or characters. Timeseries data is another: its values are organized over a sequence of positions. Other ordered inputs can also be represented as sequences before being processed by a deep-learning model.
| Input type | How it is viewed | Initial representation step |
|---|---|---|
| Text | A sequence of words or characters | Tokenize, then vectorize |
| Timeseries | Values arranged over sequence positions | Create a numeric sequence representation |
| Other ordered input | A sequence whose positions have meaning | Represent it as a numeric tensor |
Examples of ordered inputs that can be prepared for sequence processing
Mistakes to Avoid
Treating tokenization and vectorization as the same operation.
Tokenization identifies the units, while vectorization supplies numeric representations for those units.
Fix:
Describe the stages separately: first choose tokens, then turn them into vectors and pack those vectors into a sequence tensor.Feeding raw text directly into a neural network.
The source states that neural networks receive numeric tensors rather than raw text.
Fix:
Create a numerical representation before applying a deep-learning model.Calling a token embedding a sequence-processing model.
The embedding supplies vectors for tokens; the 1D convnet is the model that applies one-dimensional convolution to sequence data.
Fix:
Use “token embedding” for the representation and “1D convnet” for the sequence-processing approach.Ignoring order after vectorization.
Sequence data is defined by meaningful order, and the vectors are packed into sequence tensors.
Fix:
Preserve the sequence positions when assembling the numeric representation.
Practice Check
A learner says: “I tokenized the sentence, so it is ready for a 1D convnet.” Explain what step is still needed and why.
Hints
- Ask what tokenization identifies.
- Ask what a neural network requires as input.
- Name the operation that supplies each token with a numeric vector.
Classify each item as a sequence representation step or a sequence-processing model: tokenization, token embedding, 1D convnet, recurrent neural network.
Hints
- Representation steps prepare the numeric tensor.
- Model approaches operate on the prepared sequence.
- Tokenization identifies units; token embedding supplies vectors.
Key Takeaways
- Sequence data is ordered input, including text, timeseries, and other ordered forms.
- Neural networks process numeric tensors rather than raw text.
- Tokenization chooses the sequence units, while vectorization gives those units numeric representations.
- Token embeddings associate tokens with vectors, which are packed into sequence tensors.
- A 1D convnet is one fundamental deep-learning approach for sequence processing; a recurrent neural network is the other broad approach identified in the source.
Key Takeaways
- Ordered sequence data must be converted into numeric tensors before a neural network can process it.
- Tokenization identifies units such as words or characters; vectorization turns those units into numeric representations.
- Token embeddings associate tokens with vectors and preserve those vectors in sequence order.
- 1D convnets apply one-dimensional convolution to sequence data, while recurrent neural networks represent another broad sequence-processing approach.
- Text, timeseries, and other ordered inputs are natural settings for sequence-processing models.