Text Preprocessing and Word Embeddings
Word embeddings turn words into dense vector representations.
From Words to Numerical Representations
A neural network cannot work directly with the ordinary meaning of a word as it appears in text. A word must receive a numerical representation. Word embeddings provide that representation by turning words into dense vectors. The important question is not merely whether every word has numbers attached to it. The important question is whether those numbers preserve useful relationships among words.
A word embedding is a dense vector representation of a word. Its value comes from a learned representation intended to express meaningful relationships between words, rather than from an arbitrary assignment of numbers.
Why Random Vectors Fail
What do you think happens?
Suppose every word receives a dense vector chosen at random. Does that automatically create a useful representation space?
Reveal answer
Answer: No, because the vectors may not express meaningful relationships.
Random vectors give words numerical forms, but they do not organize words according to meaningful relationships. Two words with similar uses can receive unrelated vectors, leaving the representation space without useful structure.
Random assignment appears to solve one narrow problem: every word has a numerical form. It fails at the deeper problem of representation. If the vector for one word is selected independently of the vectors for other words, similar uses do not necessarily produce similar positions. The resulting space can place related words far apart and gives a deep neural network no reliable organization to interpret.
Numerical form is not the same as useful representation. A representation becomes useful when relationships among words are organized in a way that reflects meaningful relationships between those words.
Similarity in the Embedding Space
An embedding space can be understood as an arrangement of representational positions. Related words should occupy similar positions in that space. The source gives interchangeable words as the key requirement: when two words can be used in most of the same sentences, their embeddings should reflect that relationship. If those words are placed far apart in meaning-relevant terms, the space becomes noisy and more difficult for a deep neural network to interpret.
Interchangeable Words
Consider two words that are interchangeable in most sentences. What relationship should their embeddings express?
Identify the word relationship: The words have similar uses because they can replace one another in most sentences.
Apply the representation requirement: Their embeddings should occupy similar representational positions rather than unrelated positions.
Compare with random assignment: Random vectors could place the words far apart, which would fail to express their meaningful relationship.
Interchangeable words should receive similar representations so that the embedding space reflects their relationship.
The Embedding Layer
An embedding layer is used to learn structured word embeddings. Its purpose is not merely to attach some vector to each word. Its purpose is to learn representations whose relationships reflect meaningful relationships between words. This distinguishes an embedding layer from a random lookup: the layer supports learning an organized representation space instead of relying on arbitrary associations.
The central mechanism is learning structure. At the beginning, the goal is not satisfied by giving every word any vector. The embedding layer is valuable because it provides a way to learn word embeddings whose positions and relationships are useful. In that learned space, related or interchangeable words are expected to have similar representations instead of being placed arbitrarily far apart.
Common Representation Mistakes
Assuming that every numerical representation is automatically meaningful.
A random vector gives a word numerical form but does not organize words according to meaningful relationships.
Fix:
Ask whether the relationships among the vectors reflect relationships among the words.Treating interchangeable words as unrelated because they have different spellings.
Their similar uses should be reflected by similar representational positions.
Fix:
Expect related or interchangeable words to receive similar representations in a structured embedding space.Thinking that a structured embedding space gives every word the same vector.
Structure concerns the relationships among word representations, not the elimination of all differences.
Fix:
Keep the focus on useful organization among representations.Describing the embedding layer as only a device for attaching arbitrary vectors.
The embedding layer provides a way to learn structured word embeddings.
Fix:
Describe its role as learning representations whose relationships reflect meaningful relationships between words.
Check Your Understanding
A model assigns dense vectors to every word, but the vectors are selected independently at random. Explain why this is inadequate. Then describe what should be different when an embedding layer learns the representations.
Hints
- Separate the existence of a numerical form from the usefulness of the relationships among representations.
- Use interchangeable words as your test case.
- Explain why similar representational positions are preferable for words with similar uses.
A Strong Answer
Explain the difference between random word vectors and learned word embeddings.
Describe random vectors: Random vectors give words numerical forms, but they may place words with similar uses at unrelated positions.
Describe learned embeddings: An embedding layer learns representations so that relationships among words can be reflected in relationships among their vectors.
State the practical consequence: A structured space is less noisy and more useful for a deep neural network to interpret.
Random vectors represent words numerically but without reliable organization; learned embeddings aim to represent meaningful relationships between words.
Key Takeaways
- Word embeddings turn words into dense vector representations.
- Random vectors are inadequate because they do not organize words according to meaningful relationships.
- Related or interchangeable words should occupy similar representational positions.
- An embedding layer provides a way to learn structured word embeddings.
- A structured space represents useful relationships among words without requiring every word to have the same vector.
Key Takeaways
- Word embeddings convert words into dense vectors.
- Randomly assigned vectors provide numbers but not meaningful organization.
- Similar uses, especially interchangeability, should be reflected by similar representations.
- An embedding layer learns structured relationships among word representations.
- Useful structure means meaningful relationships among vectors, not identical vectors for every word.