Word Embeddings
TensorBoard is a browser-based visualization tool packaged with TensorFlow for inspecting Keras models using the TensorFlow backend.
From Training Run to Evidence
A final loss value tells you how one training run ended, but it does not show everything that happened inside the model. During iterative development, each experiment should provide information that helps shape the next experiment. TensorBoard supports this process by turning information recorded during TensorFlow training into browser-based visualizations.
TensorBoard is a browser-based visualization tool packaged with TensorFlow for inspecting Keras models using the TensorFlow backend.
The important handoff is simple: training produces log events, and TensorBoard reads those events to present several visual views. TensorBoard does not replace the experiment. It gives you a way to inspect the experiment while it is running or after its events have been written.
What the Dashboard Reveals
TensorBoard provides different views because different questions require different evidence. Metrics show how training and validation measurements develop over time. Histograms show distributions of activation values taken by layers, and the source also identifies gradients as information that can be viewed. The Embeddings view shows spatial relationships learned by an embedding layer after the high-dimensional space is reduced to two or three dimensions.
| TensorBoard view | What it helps you inspect | Why it is different from final loss |
|---|---|---|
| Metrics | Training and validation measurements over time | Shows the trajectory of the run |
| Histograms | Activation values and gradients | Shows internal numerical distributions |
| Embeddings | Spatial relationships among learned representations | Shows organization inside the representation space |
| Graphs | Low-level TensorFlow operations | Shows computation structure rather than a score |
For example, a metric compresses a run into a small number of measurements. A histogram provides a different perspective by showing how activation values or gradients are distributed across training. An embedding visualization can show whether learned word representations form visible spatial relationships. These views help you treat the model as more than a black box that produces only an output score.
Two Kinds of Model Graph
TensorBoard's Graphs tab and Keras's plot_model utility show related but different representations. TensorBoard shows the graph of low-level TensorFlow operations underlying a Keras model. This graph can be much more complicated than the short sequence of layers used to define the model, partly because it includes structures associated with gradient descent.
Keras provides keras.utils.plot_model when you want a cleaner layer-level picture. That utility can write a PNG image of the model graph, and passing show_shapes=True adds shape information to that layer graph. The two views are therefore not competing descriptions: one exposes lower-level operations, while the other emphasizes the layer structure used to describe the model.
A Word as a Numerical Representation
A text model cannot work directly with a word as a piece of written language. It needs a numerical representation. One-hot encoding represents each vocabulary word with a sparse binary vector: there is one dimension for each vocabulary word, one active position, and many zero positions.
A word embedding is a dense, low-dimensional floating-point vector learned from data.
The important change is not merely replacing one kind of number with another. One-hot encoding gives every vocabulary word its own dimension. An embedding instead stores each word as a compact collection of learned numeric features. The embedding dimensionality is separate from vocabulary size; the source gives 128 dimensions in its TensorBoard example and identifies 256, 512, and 1,024 as common embedding sizes.
The main advantage discussed here is dimensionality. If the vocabulary contains many words, one-hot encoding needs one dimension for every word. An embedding can use far fewer dimensions than that vocabulary size, so more information is packed into a smaller representation.
Learning Positions from Data
An embedding is not simply a shorter one-hot vector. Its numeric values are learned from data. During training, the model learns representations for words in the input vocabulary as part of the model's task. The resulting locations in the embedding space reflect relationships shaped by that training objective.
From sentiment data to word relationships
What can a learned embedding reveal after it has been trained for a sentiment task?
Represent the input: The text model begins with an embedding layer that learns representations for words in the input vocabulary.
Train for the task: The representations are learned jointly with the particular objective, so the task influences how words are positioned.
Reduce the space for inspection: TensorBoard reduces the high-dimensional embedding space to two or three dimensions using a selected method such as PCA or t-SNE.
Inspect the relationships: In the source visualization, words with positive connotations and words with negative connotations form two visible clusters.
The visualization can reveal task-shaped spatial relationships that are not visible in a one-hot representation or a final score alone.
Inspecting a Sentiment Model
The source example uses a one-dimensional convolutional network for IMDB sentiment analysis. It limits the vocabulary to the top 2,000 words, pads texts to a maximum length of 500, begins with an embedding layer, continues through convolution and pooling layers, and ends with a dense output layer.
The training run uses a TensorBoard callback to write log events to a directory. After training begins, the TensorBoard command-line utility reads those logs, and the source example opens TensorBoard at http://localhost:6006. The useful sequence is not the address itself; it is the movement from training, to recorded events, to inspectable views.
For this model, the Embeddings view can make the learned word relationships easier to inspect. Because the embedding space has 128 dimensions, TensorBoard reduces it to two or three dimensions using a selected dimensionality-reduction method such as PCA or t-SNE. The displayed clusters are an interpretation of the learned representation, not a replacement for the training objective.
Common Interpretation Mistakes
Treating a word embedding as a one-hot vector with fewer positions.
The two representations differ in both dimensional structure and how their values are obtained.
Fix:
Remember that embedding values are learned from data and can encode task-shaped relationships.Assuming embedding dimensionality must equal vocabulary size.
Vocabulary size and embedding dimensionality are separate quantities.
Fix:
Compare one-hot length with the selected embedding-vector length.Reading TensorBoard's Graphs tab as a simple layer diagram.
A low-level operation graph may be much more complicated than the layer sequence used to define the Keras model.
Fix:
Use keras.utils.plot_model for a cleaner layer-level picture, optionally with shape information.Using only final loss to judge what happened during training.
The final score compresses the run and hides internal numerical behavior.
Fix:
Use metrics, histograms, and embedding views as complementary evidence.Assuming a visible embedding cluster is universally meaningful.
The embedding was learned jointly with a particular objective, so its organization is shaped by that task.
Fix:
Interpret the visualization in relation to the data and objective used to learn it.
Check Your Understanding
A vocabulary contains 2,000 words, and a model uses a 128-dimensional embedding space. Explain how the one-hot representation and embedding representation differ in length, density, and origin of their numeric values. Then explain which TensorBoard view could help you inspect relationships among the learned word representations.
Hints
- One-hot encoding uses one dimension for each vocabulary word.
- A word embedding is dense, low-dimensional, and learned from data.
- The Embeddings view shows spatial relationships after reducing the high-dimensional space to two or three dimensions.
You want to know whether a model's internal numerical behavior is changing during training, rather than looking only at its final loss. Which TensorBoard views would you inspect, and what would each contribute?
Hints
- Metrics show measurements over time.
- Histograms show activation values and gradients.
- Embeddings show spatial relationships among learned representations.
Key Takeaways
- TensorBoard turns recorded TensorFlow training events into browser-based views that support iterative experimentation.
- Metrics show training and validation trajectories, histograms show activation values and gradients, and embedding views show spatial relationships.
- TensorBoard's Graphs tab displays low-level TensorFlow operations, while Keras plot_model provides a cleaner layer-level graph.
- One-hot encoding uses a sparse binary vector with one dimension per vocabulary word; a word embedding is a dense, low-dimensional floating-point vector.
- Embedding values are learned from data, and the resulting relationships are shaped by the task used to train the model.
Key Takeaways
- A word embedding is a dense, low-dimensional floating-point vector learned from data.
- Unlike one-hot encoding, an embedding uses a compact number of dimensions that is separate from vocabulary size.
- TensorBoard helps inspect training trajectories, internal activation and gradient distributions, and learned embedding relationships.
- TensorBoard's operation graph and Keras's layer graph provide different levels of model structure.
- Embedding visualizations must be interpreted in relation to the objective that shaped the learned representations.