Sentiment Analysis
Sequence data must be represented numerically before a deep-learning model can process it.
From Words to Signals
Sentiment analysis asks a model to process text as sequence data. The central challenge appears before the model performs any analysis: a neural network cannot receive raw text directly. The text must first be represented numerically. A useful mental model is a pipeline: text is divided into tokens, each token receives a numeric representation, and those representations are packed into sequence tensors that a deep neural network can process.
Start by asking what sequence the model will receive. Convolution comes after tokenization and numeric representation, not before.
A Worked Representation
Representing a Short Review
Trace an illustrative text sequence through tokenization and numeric representation.
Choose the sequence units: For this generated example, treat the review as a sequence of words: bright, story, and ending. Text could instead be treated as characters or as overlapping groups of consecutive words or characters called n-grams.
Tokenize the text: Tokenization splits the text into the selected units. The resulting sequence is [bright, story, ending].
Associate numeric representations: For illustration, assign token indices [12, 4, 19]. These numbers are placeholders showing that tokens must receive numeric representations; the source does not prescribe these particular values.
Pack the sequence: The token representations are kept in sequence order and packed into sequence data that can be supplied to a deep neural network.
The model does not receive the characters in the original review directly. It receives an ordered numeric representation of the selected tokens.
This example separates two decisions that are easy to confuse. Tokenization determines what the sequence elements are. Numeric representation determines how those elements are expressed for the network. If the text is tokenized as words, the sequence elements differ from a character-based representation or an n-gram representation. A one-dimensional convnet then processes whichever ordered numeric sequence the preparation stage produced.
The Convolutional Scan
A one-dimensional convnet is the one-dimensional counterpart of a two-dimensional convnet. The key difference is the shape of the data being examined. Instead of moving across two spatial dimensions, the convolution moves along one ordered sequence dimension. At each position, a filter considers a local window of sequence values and contributes to an output sequence.
Imagine the numeric sequence laid out from left to right. The filter first examines one local region, then moves along the sequence and examines the next region. Each position therefore produces information about a nearby part of the input. The resulting features remain associated with progression through the sequence rather than being created from the entire sequence in one undifferentiated step.
Tracing Nearby Text Regions
Suppose a token sequence has six ordered positions and a filter examines three positions at a time. Describe the scan without calculating numeric filter values.
First region: The filter examines positions 1 through 3 and contributes a feature for that local region.
Next region: The filter moves along the sequence and examines the next local region, positions 2 through 4.
Continue the progression: The same scanning idea continues through later nearby regions, producing further positions in the output sequence.
The convolution transforms an ordered numeric sequence into an output sequence whose features come from successive local regions.
Convolution and Recurrence
| Approach | What it processes | Central operation |
|---|---|---|
| One-dimensional convnet | An ordered numeric sequence | Examines local regions while moving along one sequence dimension |
| Recurrent neural network | An ordered sequence | Processes sequence information through a recurrent approach |
The important distinction is not that one method handles sequences and the other does not. Both are fundamental approaches to sequence processing. A one-dimensional convnet moves a filter across local regions of one ordered dimension. A recurrent neural network uses a recurrent approach to process information across a sequence. The practical choice depends on the task and model design, but both approaches share the same first implementation question: what numeric sequence representation will the model receive?
Preparing Reliable Inputs
Sending raw text directly to the neural network
A neural network cannot receive raw text directly.
Fix:
Tokenize the text, associate vectors with the tokens, and pack those vectors into sequence tensors.Changing tokenization without recognizing that the input sequence has changed
Tokenization determines what the sequence elements are.
Fix:
Choose the sequence units deliberately and keep the numeric preparation consistent with that choice.Stopping after tokenization
Tokens still need numeric representations before a deep-learning model can process them.
Fix:
Use a numeric representation such as one-hot encoding or token embedding, then pack the resulting vectors into sequence tensors.Discarding sequence order while packing the input
A one-dimensional convolution moves along one ordered sequence dimension and examines local regions in that order.
Fix:
Preserve the selected token order when creating the sequence tensor.
Check Your Mental Model
A review is first split into characters instead of words. The characters are then given numeric representations and packed in their original order. Before a one-dimensional convolution is applied, what has changed compared with a word-based representation, and what does the convolution still do?
Hints
- Focus first on what tokenization decides.
- Then identify what remains unchanged about the convolution's movement.
Practice Answer
Explain the effect of switching from word tokens to character tokens.
Changed input units: The sequence elements change because tokenization now selects characters rather than words.
Changed numeric sequence: The numeric representations and the resulting sequence tensor describe the new character sequence.
Same convolutional idea: The one-dimensional convolution still moves along one ordered sequence dimension, examining successive local regions of the numeric input.
Tokenization changes what the convnet receives, while the convolution continues to scan local regions along the resulting ordered sequence.
Key Takeaways
- Raw text must be transformed into numeric sequence data before a deep-learning model can process it.
- Tokenization chooses the sequence units, such as words, characters, or n-grams.
- Token representations are packed in order into sequence tensors.
- A one-dimensional convnet moves along one ordered sequence dimension and examines local regions.
- One-dimensional convnets and recurrent neural networks are both fundamental sequence-processing approaches, but they organize sequence processing differently.