Concepts / Tokenization

Tokenization

Sequence models operate on ordered numeric representations, not raw text.

  • Programming

From Words to Model Input

A sequence is data whose elements have an order. Text can be treated as a sequence of words or characters, just as a timeseries is another kind of ordered data. However, a deep-learning system does not receive raw words or characters directly. It receives numeric tensors arranged as sequences. Tokenization is part of the transformation that makes this possible.

select unitsassign numeric valuesarrange as a sequencesupply inputRaw textTokenswords, characters, orn-gramsNumeric vectorsSequence tensorordered numericrepresentationSequence model
What happens to raw text as it moves through tokenization and into a form that a neural network can process?

Tokenization and Vectorization

Tokenization and text vectorization are related but different operations. Tokenization selects the units that make up the sequence. Those units may be words, characters, or n-grams. Text vectorization then assigns numeric vectors to the selected units and forms sequence tensors. In short, tokenization answers which text units are present, while vectorization answers how those units are represented numerically for a model.

OperationMain questionResult
TokenizationWhich units should form the sequence?Words, characters, or n-grams
Text vectorizationHow should those units be represented numerically?Numeric vectors arranged into sequence tensors

The source distinguishes selecting sequence units from assigning their numeric representations.

producesformsTokenizationselect text unitsText vectorizationassign numeric vectorsText unitswords, characters, orn-gramsSequence tensorordered numericrepresentation
What is the difference between splitting text into tokens and converting those tokens into numeric representations?

A Worked Text Transformation

Tracing a Short Ordered Text

Trace the conceptual transformation of the text "deep models learn" when the text is treated as a sequence of words.

Select the units: Treat the text as a word sequence. The selected units are "deep", "models", and "learn". This is the tokenization step.

Retain the order: The units remain arranged in their sequence order: "deep" comes before "models", and "models" comes before "learn".

Assign numeric representations: Text vectorization assigns numeric vectors to the selected units.

Form the model input: The numeric representations are arranged as a sequence tensor, which is the kind of ordered numeric input a sequence-processing neural network can receive.

The text has moved conceptually from ordered words to an ordered numeric representation. Tokenization selected the words; vectorization supplied the numeric representation.

vector 1vector 2vector 3Position 1deepNumeric sequencevector 1, vector 2, vector3Position 2modelsPosition 3learn
How do token positions and their order remain represented after text has been converted into numbers?

Sequence Models After Vectorization

Once sequence information has been represented numerically, it can be supplied to a sequence-processing neural network. The supplied material presents recurrent neural networks and one-dimensional convolutional networks, or 1D convnets, as two fundamental deep-learning algorithms for sequence processing. A 1D convnet is described as the one-dimensional version of the 2D convolutional networks used for other data types.

represented byrepresented byrepresented bydeepNumeric vectorrepresentation of deepmodelsNumeric vectorrepresentation of modelslearnNumeric vectorrepresentation of learn
How is each token associated with the numeric representation used in a sequence tensor?

Common Misunderstandings

  • Treating tokenization and vectorization as the same operation.

    The source distinguishes selecting words, characters, or n-grams from assigning numeric vectors and forming sequence tensors.

    Fix: Use tokenization for selecting the sequence units and text vectorization for creating their numeric representations.

  • Assuming a neural network can process raw words or characters directly.

    The source states that deep-learning systems receive numeric tensors arranged as sequences, not raw words or characters.

    Fix: Explain the transformation from text units to numeric vectors and then to an ordered sequence tensor.

  • Ignoring sequence order after converting text to numbers.

    A sequence is defined by the order of its elements, and the model input is an ordered numeric representation.

    Fix: Track the position of each token's numeric representation within the sequence.

  • Inventing a specific 1D pooling calculation from the general description of 1D convnets.

    The supplied material does not contain enough information to calculate or diagram a specific 1D pooling transition.

    Fix: State only that 1D convnets are a fundamental sequence-processing family unless additional pooling specifications are provided.

Check Your Understanding

MEDIUM

A sentence is being prepared for a sequence-processing neural network. Describe, in order, what tokenization does, what text vectorization does, and what form the final model input takes. Then state whether the supplied material gives enough information to calculate a specific 1D pooling transition.

Hints
  • Begin by naming the possible units selected from text.
  • Distinguish selecting units from assigning numeric vectors.
  • Use the phrase ordered numeric representation when describing the final input.
  • Separate what the source establishes about 1D convnets from details it does not specify.

Practice Answer

Explain the preparation of a sentence for a sequence-processing neural network.

Tokenization: Select words, characters, or n-grams as the units of the text sequence.

Text vectorization: Assign numeric vectors to the selected units.

Sequence tensor: Arrange the numeric representations as an ordered sequence tensor.

Pooling specification: Do not calculate a specific 1D pooling transition from this material alone, because the required technical details are not provided.

The model receives an ordered numeric sequence tensor. Tokenization selects the units, vectorization supplies their numeric representations, and the supplied material does not specify enough information for a particular pooling calculation.

Key Takeaways

  1. A sequence is data whose elements have an order, and text can be treated as a sequence of words or characters.
  2. Tokenization selects words, characters, or n-grams; text vectorization assigns numeric vectors and forms sequence tensors.
  3. Deep-learning systems process ordered numeric representations rather than raw words or characters.
  4. Recurrent neural networks and 1D convnets are presented as fundamental deep-learning algorithms for sequence processing.
  5. The supplied material does not specify enough technical detail to calculate or diagram a particular 1D pooling transition.

Key Takeaways

  • Tokenization selects the units of a text sequence, such as words, characters, or n-grams.
  • Text vectorization assigns numeric vectors to those units and forms ordered sequence tensors.
  • Sequence order remains part of the numeric representation supplied to a sequence-processing model.
  • The source presents recurrent neural networks and 1D convnets as fundamental sequence-processing model families.
  • Specific 1D pooling behavior cannot be calculated from the supplied material without additional technical specifications.