Concepts / One-Hot Encoding

One-Hot Encoding

A word embedding is a dense, low-dimensional floating-point vector learned from data.

  • Programming

From Written Token to Numbers

A text-processing system cannot work directly with a word or character as a piece of written language. It needs a numerical representation. One-hot encoding provides a basic method: it turns a token into a long binary vector containing one active position and many zero positions.

The process has two stages. First, assign each token a unique integer index within a vocabulary. Second, use that index to create a vector with one position for each vocabulary item. The active position identifies which token is being represented.

assignedselectssuntoken3unique indexVector position 3active position
How does a token become a specific vocabulary index before that index is used to create the vector?

Assigning the Vocabulary Index

A vocabulary is the collection of tokens that a system can represent. Each token receives a unique integer index. That index is necessary because the vector must know which one of its positions represents the token. Without the index-to-position connection, a sequence of zeros and ones would not identify a particular word or character.

Finding the Active Position

Suppose a small word vocabulary assigns the token sun to index 3 and contains 5 vocabulary items. Construct the one-hot vector for sun.

Count the vector positions: The vocabulary contains 5 items, so the vector has 5 positions.

Use the assigned index: The index for sun identifies the third vector position.

Activate one position: Place 1 at the position selected by the index and place 0 in every other position.

The one-hot vector is [0, 0, 1, 0, 0]. The 1 at the third position indicates the token sun.

Reading the Binary Vector

For a vocabulary of size N, a one-hot vector has N positions and exactly one active entry. The active entry is 1, while the remaining entries are 0. The 1 does not describe a numerical property of the word. It marks the vocabulary position assigned to that token.

Position 10Position 20Position 31Position 40Position 50
What does the active 1 position indicate, and why are all other positions 0?

In the vector [0, 0, 1, 0, 0], the third position is active. Interpreting the vector requires consulting the vocabulary mapping: whichever token owns position 3 is the token represented by this vector.

Words and Characters as Tokens

The same one-hot procedure can operate at different token levels. In word-level encoding, the vocabulary items are words. In character-level encoding, the vocabulary items are characters. The chosen token level changes the vocabulary, the assigned index, and therefore the position of the active 1.

Encoding levelVocabulary itemWhat the index identifiesVector size
Word-levelA wordThe word's vocabulary positionNumber of vocabulary words
Character-levelA characterThe character's vocabulary positionNumber of vocabulary characters
maps toselectsmaps toselectsWord vocabularysunWord indexassigned positionWord vectorone active positionCharactervocabularysCharacter indexassigned positionCharacter vectorone active position
How does choosing a word or character as the token change the vocabulary, index, and resulting one-hot vector?

One-Hot Vectors and Word Embeddings

A word embedding is a dense, low-dimensional floating-point vector learned from data. One-hot encoding instead produces a sparse binary vector with one dimension for each vocabulary word. Both provide numerical representations, but they represent information in different ways.

PropertyOne-hot vectorWord embedding
ValuesBinary values: one active entry and many zero entriesFloating-point values
DensitySparseDense
DimensionsOne dimension for each vocabulary wordLow-dimensional size separate from vocabulary size
How it is createdUses a vocabulary index to select one positionLearned from data
usesusesOne-hot vectorsparse binary valuesVocabulary sizeone dimension per wordWord embeddingdense floating-point valuesEmbedding sizeseparate from vocabularysize
How do a sparse one-hot vector and a dense word-embedding vector differ in dimensionality, values, and representation?

Learning a Dense Representation

A one-hot vector is created directly from the vocabulary index. A word embedding is different: its dense floating-point vector is learned from data. Learning from data is what gives the embedding its numerical representation rather than merely marking a vocabulary position.

provides examplesproducesTraining datatext examplesLearning processuses dataWord embeddingdense floating-point vector
How does learning from data produce a dense numerical vector for a word?

This is why a word embedding should not be confused with a one-hot vector. One-hot encoding answers the question, “Which vocabulary position belongs to this token?” A word embedding is a learned numerical representation with fewer dimensions than the vocabulary-based one-hot representation.

Common Encoding Mistakes

  • Treating the token itself as the vector.

    A text-processing system needs a numerical representation, and one-hot encoding creates that representation through an index-to-position mapping.

    Fix: Map the token to its unique vocabulary index first, then activate the corresponding vector position.

  • Making the vector length unrelated to the vocabulary.

    For a vocabulary of size N, the one-hot vector has N positions.

    Fix: Use one vector position for each vocabulary item.

  • Placing more than one active entry in a one-hot vector.

    One-hot encoding uses one active position for the represented token and zero in every other position.

    Fix: Set exactly one position to 1 and all remaining positions to 0.

  • Calling a one-hot vector a word embedding.

    A one-hot vector is sparse and binary, while a word embedding is dense, low-dimensional, floating-point, and learned from data.

    Fix: Identify the representation by its values, dimensions, and creation process.

Practice the Mapping

EASY

A character vocabulary has 4 items. The character m is assigned to the second position. Write the one-hot vector for m, then explain what the active position tells you.

Hints
  • The vector needs one position for each vocabulary item.
  • The assigned position receives 1.
  • Every other position receives 0.

What do you think happens?

What should the vector look like if the token is assigned to the second position in a four-item vocabulary?

  • [1, 0, 0, 0]
  • [0, 1, 0, 0]
  • [0, 0, 1, 0]
  • [0, 0, 0, 1]
Reveal answer

Answer: [0, 1, 0, 0]

The vocabulary size determines the four positions, and the assigned index determines which one receives the single active value 1.

Key Takeaways

  1. One-hot encoding turns a token into a sparse binary vector.
  2. A unique vocabulary index connects the token to one vector position.
  3. For a vocabulary of size N, the vector has N positions and one active entry.
  4. Word-level and character-level encoding use the same procedure with different vocabulary items.
  5. A word embedding is a dense, low-dimensional floating-point vector learned from data, and it uses fewer dimensions than one-hot encoding.

Key Takeaways

  • One-hot encoding represents a token with one active binary position and zero in every other position.
  • The vocabulary index is the link between the token and the vector position.
  • The vector length equals the vocabulary size.
  • Word-level and character-level encoding differ in what counts as a vocabulary item.
  • Word embeddings are dense, low-dimensional floating-point vectors learned from data, unlike sparse one-hot vectors.