Concepts / IMDB and Reuters Text Classification Examples

IMDB and Reuters Text Classification Examples

One-hot encoding is a basic method for turning tokens into vectors.

  • Programming

From Text to Numbers

Text-processing systems cannot use a word or character directly as a vector. One-hot encoding provides a basic way to represent a token numerically. This matters when working with text such as IMDB reviews or Reuters newswire articles: before a token can be represented as a vector, the system must connect it to a position in a vocabulary.

One-hot encoding has two stages: assign a token a unique integer index, then use that index to create a binary vector.

look upreturnsselects positiontokenword or charactervocabularytoken-to-index mappinginteger indexone assigned positionbinary vectorone active entry
What happens step by step as a word or character moves from text through vocabulary lookup to its vector representation?

Why the Index Comes First

A vocabulary is a collection of the tokens being represented. Each token receives a unique integer index. That index is needed because the vector itself is organized by positions: the index tells the encoding process which position should be active.

For a vocabulary of size N, a one-hot vector has N positions and one active entry. The active entry is the position associated with the token's unique integer index.

Mapping a word to a position

Suppose a generated vocabulary assigns the word "excellent" the unique index 2 in a vocabulary with 5 tokens. Construct its one-hot vector.

Identify the vocabulary size: The vocabulary has 5 tokens, so the vector must contain 5 positions.

Identify the assigned index: The word "excellent" has been assigned index 2 in this illustrative vocabulary.

Activate that position: Place the single active entry at position 2 and leave the other positions inactive.

The illustrative one-hot vector is [0, 0, 1, 0, 0]. The third displayed entry is active because the example assigns "excellent" the index 2.

hasselectsactivatesexcellenttoken2assigned indexposition 2active position[0, 0, 1, 0, 0]five positions
How does a token's vocabulary index determine which position becomes active in its vector?

Building the One-Hot Vector

The vector length comes from the vocabulary size, not from the number of letters in a word or the number of words in a text. Once the vocabulary index is known, the corresponding vector position is made active and the remaining positions are inactive. This creates a binary vector with one active entry.

encodeexcellentindex 2[0, 0, 1, 0, 0]one active entry
What changes when a token's vocabulary index is converted into a vector with one active position?

Encoding a character

Suppose a generated character vocabulary has 4 entries and assigns the character "a" the unique index 1. Construct its one-hot vector.

Use the vocabulary size: There are 4 vocabulary entries, so the vector has 4 positions.

Use the character index: The character "a" has index 1 in this illustrative vocabulary.

Place the active entry: The entry at position 1 is active, while the other three entries are inactive.

The illustrative vector is [0, 1, 0, 0].

The active position does not describe a word's meaning. It identifies the token through the token's position in the chosen vocabulary.

Reading the Active Position

To interpret a one-hot vector, locate its single active entry and then consult the same vocabulary used to create the vector. The position points back to the token assigned to that position. Without the vocabulary, the vector's active position is only a position; the vocabulary supplies the token label.

Decoding an active position

A generated vocabulary records the token "news" at position 3. Which token is represented by the vector [0, 0, 0, 1]?

Locate the active entry: The only active entry appears at position 3 in the illustrative vector.

Check the vocabulary: The generated vocabulary records "news" at position 3.

The vector represents "news" within that generated vocabulary.

represented byposition 1reviewposition 2articleposition 3news[0, 0, 0, 1]active position 3
How can you identify which token a one-hot vector represents by locating its active position?

Words and Characters

AspectWord-level encodingCharacter-level encoding
Vocabulary itemsWordsCharacters
Token boundaryEach token is a wordEach token is a character
Vector lengthThe size of the word vocabularyThe size of the character vocabulary
Encoding procedureAssign each word a unique index, then activate its positionAssign each character a unique index, then activate its position

Word-level and character-level encoding differ in what counts as a vocabulary item. With word-level encoding, the vocabulary contains words. With character-level encoding, it contains characters. The procedure remains the same: assign a unique integer index to the token and use that index to create a vector with one active entry.

containsencodecontainsencodeword vocabularyreview, article, newscharactervocabularya, r, tarticleword tokenacharacter tokenword vectorlength equals wordvocabulary sizecharacter vectorlength equals charactervocabulary size
How do the vocabulary, token boundaries, vector length, and active positions differ between words and characters?

Mistakes with One-Hot Vectors

  • Skipping the vocabulary index

    The index identifies which vector position should be active.

    Fix: Record the token's vocabulary index before constructing the vector.

  • Making the vector length depend on the token's spelling

    For a vocabulary of size N, the vector has N positions.

    Fix: Use the size of the relevant vocabulary to determine the vector length.

  • Treating a word and a character as the same vocabulary item

    Word-level and character-level encoding use different vocabulary items.

    Fix: Decide whether the token is a word or a character, then use the matching vocabulary.

  • Reading the active position without the vocabulary

    The position identifies the token only through the vocabulary that assigned it.

    Fix: Interpret the active position by looking it up in the same vocabulary used during encoding.

Practice and Summary

EASY

A generated vocabulary has 6 entries. The word "review" is assigned index 4. Construct its one-hot vector, then explain how you would identify the represented word when given that vector.

Hints
  • Start with the vocabulary size to determine the number of vector positions.
  • Use the assigned index to place the single active entry.
  • To interpret the vector, use the same vocabulary to look up the active position.
  1. One-hot encoding turns a token into a binary vector with one active entry. The process first assigns the token a unique integer index, then uses that index to select the active vector position. For a vocabulary of size N, the vector has N positions. The active position can be interpreted only with the vocabulary that assigned the index. Word-level and character-level encoding follow the same procedure, but they use different vocabulary items.

Key Takeaways

  • One-hot encoding provides a basic numerical representation for a word or character.
  • A unique vocabulary index is required before the vector can be constructed.
  • A vocabulary of size N produces a vector with N positions and one active entry.
  • The active position identifies the token through the vocabulary used for encoding.
  • Word-level and character-level encoding use the same procedure with different vocabulary items.