IMDB and Reuters Text Classification Examples
One-hot encoding is a basic method for turning tokens into vectors.
From Text to Numbers
Text-processing systems cannot use a word or character directly as a vector. One-hot encoding provides a basic way to represent a token numerically. This matters when working with text such as IMDB reviews or Reuters newswire articles: before a token can be represented as a vector, the system must connect it to a position in a vocabulary.
One-hot encoding has two stages: assign a token a unique integer index, then use that index to create a binary vector.
Why the Index Comes First
A vocabulary is a collection of the tokens being represented. Each token receives a unique integer index. That index is needed because the vector itself is organized by positions: the index tells the encoding process which position should be active.
For a vocabulary of size N, a one-hot vector has N positions and one active entry. The active entry is the position associated with the token's unique integer index.
Mapping a word to a position
Suppose a generated vocabulary assigns the word "excellent" the unique index 2 in a vocabulary with 5 tokens. Construct its one-hot vector.
Identify the vocabulary size: The vocabulary has 5 tokens, so the vector must contain 5 positions.
Identify the assigned index: The word "excellent" has been assigned index 2 in this illustrative vocabulary.
Activate that position: Place the single active entry at position 2 and leave the other positions inactive.
The illustrative one-hot vector is [0, 0, 1, 0, 0]. The third displayed entry is active because the example assigns "excellent" the index 2.
Building the One-Hot Vector
The vector length comes from the vocabulary size, not from the number of letters in a word or the number of words in a text. Once the vocabulary index is known, the corresponding vector position is made active and the remaining positions are inactive. This creates a binary vector with one active entry.
Encoding a character
Suppose a generated character vocabulary has 4 entries and assigns the character "a" the unique index 1. Construct its one-hot vector.
Use the vocabulary size: There are 4 vocabulary entries, so the vector has 4 positions.
Use the character index: The character "a" has index 1 in this illustrative vocabulary.
Place the active entry: The entry at position 1 is active, while the other three entries are inactive.
The illustrative vector is [0, 1, 0, 0].
The active position does not describe a word's meaning. It identifies the token through the token's position in the chosen vocabulary.
Reading the Active Position
To interpret a one-hot vector, locate its single active entry and then consult the same vocabulary used to create the vector. The position points back to the token assigned to that position. Without the vocabulary, the vector's active position is only a position; the vocabulary supplies the token label.
Decoding an active position
A generated vocabulary records the token "news" at position 3. Which token is represented by the vector [0, 0, 0, 1]?
Locate the active entry: The only active entry appears at position 3 in the illustrative vector.
Check the vocabulary: The generated vocabulary records "news" at position 3.
The vector represents "news" within that generated vocabulary.
Words and Characters
| Aspect | Word-level encoding | Character-level encoding |
|---|---|---|
| Vocabulary items | Words | Characters |
| Token boundary | Each token is a word | Each token is a character |
| Vector length | The size of the word vocabulary | The size of the character vocabulary |
| Encoding procedure | Assign each word a unique index, then activate its position | Assign each character a unique index, then activate its position |
Word-level and character-level encoding differ in what counts as a vocabulary item. With word-level encoding, the vocabulary contains words. With character-level encoding, it contains characters. The procedure remains the same: assign a unique integer index to the token and use that index to create a vector with one active entry.
Mistakes with One-Hot Vectors
Skipping the vocabulary index
The index identifies which vector position should be active.
Fix:
Record the token's vocabulary index before constructing the vector.Making the vector length depend on the token's spelling
For a vocabulary of size N, the vector has N positions.
Fix:
Use the size of the relevant vocabulary to determine the vector length.Treating a word and a character as the same vocabulary item
Word-level and character-level encoding use different vocabulary items.
Fix:
Decide whether the token is a word or a character, then use the matching vocabulary.Reading the active position without the vocabulary
The position identifies the token only through the vocabulary that assigned it.
Fix:
Interpret the active position by looking it up in the same vocabulary used during encoding.
Practice and Summary
A generated vocabulary has 6 entries. The word "review" is assigned index 4. Construct its one-hot vector, then explain how you would identify the represented word when given that vector.
Hints
- Start with the vocabulary size to determine the number of vector positions.
- Use the assigned index to place the single active entry.
- To interpret the vector, use the same vocabulary to look up the active position.
- One-hot encoding turns a token into a binary vector with one active entry. The process first assigns the token a unique integer index, then uses that index to select the active vector position. For a vocabulary of size N, the vector has N positions. The active position can be interpreted only with the vocabulary that assigned the index. Word-level and character-level encoding follow the same procedure, but they use different vocabulary items.
Key Takeaways
- One-hot encoding provides a basic numerical representation for a word or character.
- A unique vocabulary index is required before the vector can be constructed.
- A vocabulary of size N produces a vector with N positions and one active entry.
- The active position identifies the token through the vocabulary used for encoding.
- Word-level and character-level encoding use the same procedure with different vocabulary items.