Concepts / Sentiment Analysis

Sentiment Analysis

Sequence data must be represented numerically before a deep-learning model can process it.

  • Programming

From Words to Signals

Sentiment analysis asks a model to process text as sequence data. The central challenge appears before the model performs any analysis: a neural network cannot receive raw text directly. The text must first be represented numerically. A useful mental model is a pipeline: text is divided into tokens, each token receives a numeric representation, and those representations are packed into sequence tensors that a deep neural network can process.

Start by asking what sequence the model will receive. Convolution comes after tokenization and numeric representation, not before.

tokenizationassociate vectorspack in orderRaw texta sentenceTokenswords, characters, orn-gramsToken vectorsnumeric representationsSequence tensormodel input
How does raw text become tokens, numeric representations, and a sequence tensor?

A Worked Representation

Representing a Short Review

Trace an illustrative text sequence through tokenization and numeric representation.

Choose the sequence units: For this generated example, treat the review as a sequence of words: bright, story, and ending. Text could instead be treated as characters or as overlapping groups of consecutive words or characters called n-grams.

Tokenize the text: Tokenization splits the text into the selected units. The resulting sequence is [bright, story, ending].

Associate numeric representations: For illustration, assign token indices [12, 4, 19]. These numbers are placeholders showing that tokens must receive numeric representations; the source does not prescribe these particular values.

Pack the sequence: The token representations are kept in sequence order and packed into sequence data that can be supplied to a deep neural network.

The model does not receive the characters in the original review directly. It receives an ordered numeric representation of the selected tokens.

This example separates two decisions that are easy to confuse. Tokenization determines what the sequence elements are. Numeric representation determines how those elements are expressed for the network. If the text is tokenized as words, the sequence elements differ from a character-based representation or an n-gram representation. A one-dimensional convnet then processes whichever ordered numeric sequence the preparation stage produced.

The Convolutional Scan

A one-dimensional convnet is the one-dimensional counterpart of a two-dimensional convnet. The key difference is the shape of the data being examined. Instead of moving across two spatial dimensions, the convolution moves along one ordered sequence dimension. At each position, a filter considers a local window of sequence values and contributes to an output sequence.

Imagine the numeric sequence laid out from left to right. The filter first examines one local region, then moves along the sequence and examines the next region. Each position therefore produces information about a nearby part of the input. The resulting features remain associated with progression through the sequence rather than being created from the entire sequence in one undifferentiated step.

select regionfilter responsemove along orderfilter responseOrdered sequencenumeric valuesLocal window 1first nearby regionFeature 1output at one positionLocal window 2next nearby regionFeature 2output at next position
How does a one-dimensional convolution examine successive local regions and produce features along the sequence?

Tracing Nearby Text Regions

Suppose a token sequence has six ordered positions and a filter examines three positions at a time. Describe the scan without calculating numeric filter values.

First region: The filter examines positions 1 through 3 and contributes a feature for that local region.

Next region: The filter moves along the sequence and examines the next local region, positions 2 through 4.

Continue the progression: The same scanning idea continues through later nearby regions, producing further positions in the output sequence.

The convolution transforms an ordered numeric sequence into an output sequence whose features come from successive local regions.

Convolution and Recurrence

ApproachWhat it processesCentral operation
One-dimensional convnetAn ordered numeric sequenceExamines local regions while moving along one sequence dimension
Recurrent neural networkAn ordered sequenceProcesses sequence information through a recurrent approach

The important distinction is not that one method handles sequences and the other does not. Both are fundamental approaches to sequence processing. A one-dimensional convnet moves a filter across local regions of one ordered dimension. A recurrent neural network uses a recurrent approach to process information across a sequence. The practical choice depends on the task and model design, but both approaches share the same first implementation question: what numeric sequence representation will the model receive?

scanprocess recurrentlyNumeric sequenceordered inputNumeric sequenceordered inputLocal regionsmoving filterRecurrent processingacross the sequence
How does local scanning by a one-dimensional convnet differ from recurrent processing across an ordered sequence?

Preparing Reliable Inputs

  • Sending raw text directly to the neural network

    A neural network cannot receive raw text directly.

    Fix: Tokenize the text, associate vectors with the tokens, and pack those vectors into sequence tensors.

  • Changing tokenization without recognizing that the input sequence has changed

    Tokenization determines what the sequence elements are.

    Fix: Choose the sequence units deliberately and keep the numeric preparation consistent with that choice.

  • Stopping after tokenization

    Tokens still need numeric representations before a deep-learning model can process them.

    Fix: Use a numeric representation such as one-hot encoding or token embedding, then pack the resulting vectors into sequence tensors.

  • Discarding sequence order while packing the input

    A one-dimensional convolution moves along one ordered sequence dimension and examines local regions in that order.

    Fix: Preserve the selected token order when creating the sequence tensor.

tokenizevectorize and packprocessRaw textsentenceSequence tensornumeric ordered inputTokensselected unitsLocal-region scanone sequence dimension
What changes when text is prepared correctly for a sequence model, and what remains incomplete when preprocessing stops too early?

Check Your Mental Model

MEDIUM

A review is first split into characters instead of words. The characters are then given numeric representations and packed in their original order. Before a one-dimensional convolution is applied, what has changed compared with a word-based representation, and what does the convolution still do?

Hints
  • Focus first on what tokenization decides.
  • Then identify what remains unchanged about the convolution's movement.

Practice Answer

Explain the effect of switching from word tokens to character tokens.

Changed input units: The sequence elements change because tokenization now selects characters rather than words.

Changed numeric sequence: The numeric representations and the resulting sequence tensor describe the new character sequence.

Same convolutional idea: The one-dimensional convolution still moves along one ordered sequence dimension, examining successive local regions of the numeric input.

Tokenization changes what the convnet receives, while the convolution continues to scan local regions along the resulting ordered sequence.

Key Takeaways

  • Raw text must be transformed into numeric sequence data before a deep-learning model can process it.
  • Tokenization chooses the sequence units, such as words, characters, or n-grams.
  • Token representations are packed in order into sequence tensors.
  • A one-dimensional convnet moves along one ordered sequence dimension and examines local regions.
  • One-dimensional convnets and recurrent neural networks are both fundamental sequence-processing approaches, but they organize sequence processing differently.