Text Vectorization
Sequence models operate on ordered numeric representations, not raw text.
From Words to Model Input
A sentence is meaningful to a person because its words and their order carry meaning. A deep-learning model does not receive those raw words directly. Sequence models operate on ordered numeric representations, so text must be transformed into numeric tensors before a neural network can process it.
Why Numeric Sequences Matter
A sequence is data whose elements have an order. Text can be treated as a sequence of words or characters, and a timeseries is another kind of ordered data. The important property is that the elements are not merely an unordered collection: their positions in the sequence matter.
Deep-learning systems can process sequence inputs, but they do not receive raw words or characters directly. They receive numeric tensors arranged as sequences. Text vectorization supplies the transformation from human-readable text to that numeric form.
The central idea is not simply that text becomes numbers. Text becomes an ordered numeric representation, so the sequence structure remains available to the model.
Tokenization and Vectorization
| Stage | What it does | What it produces |
|---|---|---|
| Tokenization | Selects units such as words, characters, or n-grams | A sequence of selected text units |
| Text vectorization | Assigns numeric vectors to the selected units and forms sequence tensors | An ordered numeric representation |
Tokenization and text vectorization are related but different. Tokenization selects what counts as an item in the input. Those items may be words, characters, or n-grams. Vectorization then assigns numeric vectors to the selected items and uses them to form sequence tensors.
Word-Level Representation
Tracing a Word-Based Representation
Consider the text: "Models process sequences." Apply the word-level method described in the source.
Start with raw text: The input is still one piece of human-readable text and is not yet a numeric tensor.
Segment into words: Treat the text as a sequence of individual words: "Models", "process", and "sequences". This is the tokenization stage in this example.
Transform each word: Assign a numeric vector to each resulting word. The source establishes that each word is transformed into a vector, but it does not specify the vector values or their dimensions.
Keep the order: Place the word vectors in the same sequence order as the words. Together, they form a sequence of word-level representations.
The raw sentence has been transformed into an ordered sequence of numeric word representations, suitable as a numeric sequence input for a deep-learning model.
Sequence Models After Vectorization
Once sequence information has been represented numerically, it can be supplied to a sequence-processing neural network. The source presents recurrent neural networks and one-dimensional convolutional networks as the two fundamental deep-learning algorithms for sequence processing.
A 1D convnet is described as the one-dimensional version of the 2D convolutional networks used for other data types. The supplied material establishes that a 1D convnet can receive a text sequence after vectorization. It does not provide enough information to calculate or diagram a specific pooling transition.
Mistakes in the Transformation Pipeline
Treating tokenization as the complete vectorization process
The split produces selected text units, but the source distinguishes this step from assigning numeric vectors and forming sequence tensors.
Fix:
Describe the word split as tokenization, then describe the transformation of each word into a vector as the vectorization stage.Discarding word order after assigning vectors
A sequence is defined by ordered elements, and the model receives numeric tensors arranged as sequences.
Fix:
Keep the vector for each word in the corresponding position in the word sequence.Assuming that raw text can be sent directly to a neural network
The source states that deep-learning models require numeric tensors rather than raw text.
Fix:
Describe the path from raw text through segmentation and vector conversion to an ordered numeric representation.Inventing pooling behavior from the fact that a 1D convnet is used
The provided material does not contain enough information to calculate or diagram a specific 1D pooling transition.
Fix:
State only that 1D convnets are a fundamental sequence-model family and identify pooling details as requiring additional specification.
Pipeline Practice
Describe the stages that would be needed to turn the text "Sequence models process ordered data" into an input representation for a deep-learning model using the word-level method.
Hints
- First identify the raw text.
- Next identify the selected word units.
- Then explain what happens to each word.
- Finish by explaining how the resulting vectors are arranged.
What do you think happens?
After the words have been transformed into vectors, should the vectors be treated as an unordered collection or as an ordered sequence?
Reveal answer
Answer: An ordered sequence
The source defines sequence data by the order of its elements and describes the result as numeric tensors arranged as sequences.
Key Takeaways
- Sequence models operate on ordered numeric representations, not raw text.
- Text vectorization changes raw text into numeric tensors.
- Tokenization selects words, characters, or n-grams; vectorization assigns numeric vectors and forms sequence tensors.
- One word-level method segments text into words, transforms each word into a vector, and preserves the resulting sequence order.
- The source identifies recurrent neural networks and 1D convnets as fundamental sequence-processing model families, but it does not specify enough pooling details for a concrete pooling calculation.
Key Takeaways
- Raw text must be transformed because deep-learning models require numeric tensors.
- Tokenization selects the text units, while text vectorization assigns numeric vectors and forms an ordered sequence representation.
- A word-level method moves from raw text to words, then from words to vectors, while preserving position in the sequence.
- The resulting numeric sequence can be supplied to sequence-processing models such as recurrent neural networks and 1D convnets.
- Specific 1D pooling behavior cannot be determined without additional technical details.