Concepts / Text Data Preprocessing

Text Data Preprocessing

Sequence data includes text, timeseries, and other ordered inputs.

  • Programming

From Words to Model Input

A neural network does not receive raw text directly. Before a model can process a sentence, the text must be represented as numeric data. This gives sequence processing two broad stages: create a numerical representation of the sequence, then apply a model designed to process ordered data.

tokenizationvectorizationpack into sequenceRaw textordered inputTokenswords, subwords, orcharactersNumeric vectorsone vector per tokenSequence tensorneural-network input
How does raw text move through tokenization and vectorization to become input that a neural network can process?

The important boundary is between human-readable text and numeric model input. Tokenization identifies the sequence units; vectorization gives those units numeric representations; the resulting vectors are packed into sequence tensors.

Token Units and Numeric Values

Tokenization is the step that chooses and identifies the units in a sequence. For text, those units may be words, subwords, or characters. Tokenization does not by itself create the numeric vectors needed by a neural network.

Text vectorization is the step that supplies each generated token with a numeric vector. Those vectors can then be packed into sequence tensors and passed to a deep neural network.

vectorizevectorizevectorizeTexttoken at position 0Vector 0numeric representationdatatoken at position 1Vector 1numeric representationflowstoken at position 2Vector 2numeric representation
How are sequence units identified and then associated with numeric values?

A Three-Token Sequence

Illustrate the difference between tokenization and vectorization for the generated phrase "Text data flows".

Tokenize: Treat the phrase as three word tokens: "Text", "data", and "flows". Tokenization has identified the ordered units, but the result is still text.

Vectorize: Associate each token with a numeric vector. The particular numeric values are implementation choices in this illustration; the important point is that every token receives a numeric representation.

Pack: Place the vectors in their sequence order to form a sequence tensor that can be supplied to a neural network.

Tokenization determines what the units are. Vectorization turns those units into numeric representations. Packing preserves their order in the model input.

Ways to Represent Tokens

Two major ways to associate vectors with tokens are one-hot encoding and token embedding. The source material presents token embedding as typically being used for words, where it is also called word embedding. Both belong to the vectorization stage because they provide numeric representations for the generated tokens.

Representation approachRole in preprocessingSource terminology
One-hot encodingAssociates a numeric representation with each tokenOne of the two major approaches
Token embeddingAssociates a vector with each tokenTypically used for words; also called word embedding

When describing a text pipeline, state both decisions explicitly: which units are tokens and which numeric representation is assigned to each token. This makes it clear how raw text becomes a sequence tensor.

Sliding Windows over Sequences

Once a sequence has become numeric input, a one-dimensional convnet applies the one-dimensional convolutional approach to sequence processing. Its broad role is to examine local regions of an ordered sequence as a sliding operation. This makes it suitable for detecting patterns that appear within nearby sequence positions.

inspectslideslidedetectOrdered sequencetoken positionsLocal region 1nearby positionsLocal region 2shifted positionsLocal region 3later positionsLocal patternsdetected features
How does a one-dimensional convolution examine different local regions of an ordered sequence?

The sequence remains ordered while the local region being examined moves across it. The central idea is not simply that the input is numeric; it is that the numeric input still represents an ordered sequence whose local structure can be examined.

Convolutional and Recurrent Roles

One-dimensional convnets and recurrent neural networks are presented as two fundamental deep-learning approaches for sequence processing. A one-dimensional convnet broadly scans local regions of the sequence to detect local patterns. An RNN broadly processes sequence information recurrently over time. The distinction is about the model's way of processing the ordered input, not about whether preprocessing is required: both approaches still need numeric sequence tensors rather than raw text.

scandetectprocesscarry through sequenceNumeric sequenceordered tensorNumeric sequenceordered tensorLocal scanone-dimensional convolutionRecurrent processingover timeLocal patternsdetected featuresSequence informationprocessed recurrently
What is the broad difference between a one-dimensional convnet scanning local regions and an RNN processing sequence information recurrently over time?

Where Local Patterns Matter

A one-dimensional convnet is a reasonable sequence-processing choice when the task can benefit from detecting patterns in nearby positions of an ordered input. Text and timeseries are both examples of sequence data, and other ordered inputs can be treated in the same general way. The model choice should follow the structure of the sequence and the kind of pattern the task requires the model to process.

can processcan processcan process1D convnetlocal pattern detectorTextword or character sequenceTimeseriesordered positionsOther ordered inputsequence data
Which sequence-processing inputs can be considered when the task benefits from detecting local patterns?

Consider a text task in which nearby words or characters form useful patterns. After the text has been tokenized and vectorized, a one-dimensional convnet can scan the resulting sequence to detect local features. The same broad reasoning can apply to timeseries data because timeseries are also organized over ordered positions.

Common Preprocessing Mistakes

  • Treating raw text as if it were already neural-network input

    A neural network needs a numeric tensor rather than raw text.

    Fix: State the tokenization step, the vectorization step, and the packing of vectors into a sequence tensor.

  • Calling tokenization and vectorization the same operation

    Tokenization identifies the units, while vectorization supplies each token with a numeric vector.

    Fix: Describe tokenization first and vectorization second.

  • Ignoring order in sequence data

    Text, timeseries, and other sequence data are ordered inputs.

    Fix: Preserve the sequence positions when describing the numeric tensor and the model's processing.

  • Describing a one-dimensional convnet as recurrent processing

    The broad roles differ: a one-dimensional convnet scans local regions, while an RNN processes sequence information recurrently over time.

    Fix: Use local scanning and local-pattern detection for the convnet, and recurrent processing over time for the RNN.

Practice Check

EASY

A sentence is split into word units, and each word is then associated with a numeric vector. Explain which operation performed tokenization, which performed vectorization, and what must happen before the result can be passed to a neural network.

Hints
  • Tokenization identifies the units in the sequence.
  • Vectorization supplies a numeric vector for each generated token.
  • The vectors are packed into a sequence tensor.
EASY

Choose between these descriptions: a model scans nearby positions to detect local patterns, or a model processes sequence information recurrently over time. Which description matches a one-dimensional convnet, and which matches an RNN?

Hints
  • Look for the phrase local regions.
  • Look for the phrase recurrently over time.

Key Takeaways

  • Sequence data includes text, timeseries, and other ordered inputs.
  • A neural network needs a numeric tensor rather than raw text.
  • Tokenization identifies units such as words, subwords, or characters; vectorization assigns those units numeric vectors.
  • One-hot encoding and token embedding are two major ways to associate vectors with tokens, with token embedding typically used for words and also called word embedding.
  • A one-dimensional convnet scans local regions to detect patterns, while an RNN processes sequence information recurrently over time.