One-Hot Encoding
A word embedding is a dense, low-dimensional floating-point vector learned from data.
From Written Token to Numbers
A text-processing system cannot work directly with a word or character as a piece of written language. It needs a numerical representation. One-hot encoding provides a basic method: it turns a token into a long binary vector containing one active position and many zero positions.
The process has two stages. First, assign each token a unique integer index within a vocabulary. Second, use that index to create a vector with one position for each vocabulary item. The active position identifies which token is being represented.
Assigning the Vocabulary Index
A vocabulary is the collection of tokens that a system can represent. Each token receives a unique integer index. That index is necessary because the vector must know which one of its positions represents the token. Without the index-to-position connection, a sequence of zeros and ones would not identify a particular word or character.
Finding the Active Position
Suppose a small word vocabulary assigns the token sun to index 3 and contains 5 vocabulary items. Construct the one-hot vector for sun.
Count the vector positions: The vocabulary contains 5 items, so the vector has 5 positions.
Use the assigned index: The index for sun identifies the third vector position.
Activate one position: Place 1 at the position selected by the index and place 0 in every other position.
The one-hot vector is [0, 0, 1, 0, 0]. The 1 at the third position indicates the token sun.
Reading the Binary Vector
For a vocabulary of size N, a one-hot vector has N positions and exactly one active entry. The active entry is 1, while the remaining entries are 0. The 1 does not describe a numerical property of the word. It marks the vocabulary position assigned to that token.
In the vector [0, 0, 1, 0, 0], the third position is active. Interpreting the vector requires consulting the vocabulary mapping: whichever token owns position 3 is the token represented by this vector.
Words and Characters as Tokens
The same one-hot procedure can operate at different token levels. In word-level encoding, the vocabulary items are words. In character-level encoding, the vocabulary items are characters. The chosen token level changes the vocabulary, the assigned index, and therefore the position of the active 1.
| Encoding level | Vocabulary item | What the index identifies | Vector size |
|---|---|---|---|
| Word-level | A word | The word's vocabulary position | Number of vocabulary words |
| Character-level | A character | The character's vocabulary position | Number of vocabulary characters |
One-Hot Vectors and Word Embeddings
A word embedding is a dense, low-dimensional floating-point vector learned from data. One-hot encoding instead produces a sparse binary vector with one dimension for each vocabulary word. Both provide numerical representations, but they represent information in different ways.
| Property | One-hot vector | Word embedding |
|---|---|---|
| Values | Binary values: one active entry and many zero entries | Floating-point values |
| Density | Sparse | Dense |
| Dimensions | One dimension for each vocabulary word | Low-dimensional size separate from vocabulary size |
| How it is created | Uses a vocabulary index to select one position | Learned from data |
Learning a Dense Representation
A one-hot vector is created directly from the vocabulary index. A word embedding is different: its dense floating-point vector is learned from data. Learning from data is what gives the embedding its numerical representation rather than merely marking a vocabulary position.
This is why a word embedding should not be confused with a one-hot vector. One-hot encoding answers the question, “Which vocabulary position belongs to this token?” A word embedding is a learned numerical representation with fewer dimensions than the vocabulary-based one-hot representation.
Common Encoding Mistakes
Treating the token itself as the vector.
A text-processing system needs a numerical representation, and one-hot encoding creates that representation through an index-to-position mapping.
Fix:
Map the token to its unique vocabulary index first, then activate the corresponding vector position.Making the vector length unrelated to the vocabulary.
For a vocabulary of size N, the one-hot vector has N positions.
Fix:
Use one vector position for each vocabulary item.Placing more than one active entry in a one-hot vector.
One-hot encoding uses one active position for the represented token and zero in every other position.
Fix:
Set exactly one position to 1 and all remaining positions to 0.Calling a one-hot vector a word embedding.
A one-hot vector is sparse and binary, while a word embedding is dense, low-dimensional, floating-point, and learned from data.
Fix:
Identify the representation by its values, dimensions, and creation process.
Practice the Mapping
A character vocabulary has 4 items. The character m is assigned to the second position. Write the one-hot vector for m, then explain what the active position tells you.
Hints
- The vector needs one position for each vocabulary item.
- The assigned position receives 1.
- Every other position receives 0.
What do you think happens?
What should the vector look like if the token is assigned to the second position in a four-item vocabulary?
Reveal answer
Answer: [0, 1, 0, 0]
The vocabulary size determines the four positions, and the assigned index determines which one receives the single active value 1.
Key Takeaways
- One-hot encoding turns a token into a sparse binary vector.
- A unique vocabulary index connects the token to one vector position.
- For a vocabulary of size N, the vector has N positions and one active entry.
- Word-level and character-level encoding use the same procedure with different vocabulary items.
- A word embedding is a dense, low-dimensional floating-point vector learned from data, and it uses fewer dimensions than one-hot encoding.
Key Takeaways
- One-hot encoding represents a token with one active binary position and zero in every other position.
- The vocabulary index is the link between the token and the vector position.
- The vector length equals the vocabulary size.
- Word-level and character-level encoding differ in what counts as a vocabulary item.
- Word embeddings are dense, low-dimensional floating-point vectors learned from data, unlike sparse one-hot vectors.