Concepts / Training and Validation Metrics

Training and Validation Metrics

TensorBoard is a browser-based visualization tool packaged with TensorFlow for inspecting Keras models using the TensorFlow backend.

  • Programming

From Run to Evidence

A final training score tells you where an experiment ended, but not how it got there or what happened inside the model along the way. TensorBoard helps turn a TensorFlow training run into evidence that you can inspect. It is a browser-based visualization tool packaged with TensorFlow for inspecting Keras models when Keras uses the TensorFlow backend.

writesreadspresentsTensorFlow trainingexperimentLog eventsrecorded informationTensorBoardbrowser-based viewsExperiment inspectioncompare and refine
How does information move from a running training process into visual views that a developer can inspect?

The Logging Handoff

The important operational sequence is a handoff. A training run produces log events through a TensorBoard callback. The TensorBoard command-line utility then reads those events and presents them in a browser. In the source example, TensorBoard is opened at http://localhost:6006 after training begins. Before logging, the experiment has produced information; after the events are available to TensorBoard, that information can be inspected through several visual views.

Following One Experiment

A developer is training a Keras model and wants to understand more than its final score.

Record: The training process writes log events through a TensorBoard callback.

Read: The TensorBoard command-line utility reads the recorded events.

Inspect: The developer opens TensorBoard in a browser and examines metrics, internal values, representations, or graph structure.

Refine: The observations provide information that can shape the next experiment.

TensorBoard does not replace the training experiment. It makes information from that experiment inspectable while the run is progressing or after its events have been written.

showsshowsshowsshowsshowsTensorBoardtraining inspectionMetricstraining and validationGraphslow-level operationsHistogramsactivations and gradientsEmbeddingsspatial relationshipsImagesvisual view
What categories of evidence can TensorBoard provide during model training?

Reading Metric Trajectories

The metrics view is the first place to look when you want to follow learning over time. TensorBoard provides live graphs of training and validation metrics. These graphs show how recorded measurements develop during a run, so you can inspect the trajectory rather than relying only on a single final value.

Comparing Training and Validation Records

A developer opens the metrics view and sees separate training and validation curves across the run. What should the developer learn from this view?

Identify the measurements: Determine which recorded metric each curve represents and whether it belongs to the training data or validation data.

Follow the run: Examine how each measurement develops across the training process instead of looking only at its final recorded value.

Compare the trajectories: Look at how the training and validation measurements relate over time. A difference in their trajectories is information about the experiment that a single final score would not show.

Use the observation: Treat the observed trajectory as evidence for shaping the next experiment, rather than as an automatic explanation of the model.

The metrics view changes the question from 'What was the final score?' to 'How did the recorded training and validation measurements develop during the run?'

recordsrecordsdevelops over timedevelops over timeRun beginsrecorded metricsTraining metricstep or epoch recordValidation metricstep or epoch recordLater measurementtrajectory continues
How do training and validation measurements become a trajectory rather than a single final value?

Beyond the Metrics View

Metrics compress a training run into a small number of measurements. Histograms provide a different perspective by showing distributions of activation values taken by layers. The source also identifies histograms of gradients as part of TensorBoard's feature set. These views let you examine internal numerical behavior instead of treating the model as a black box that produces only an output score.

compare stepsinspect internal behaviorActivation valuesearly training distributionGradient valueslayer distributionActivation valueslater training distribution
How can the distribution of layer activation values change across training steps, and what does that reveal beyond a metric curve?

An activation histogram is useful because it preserves distributional information. A metric may summarize a run with one value at a time, while a histogram lets you inspect the layer's activation values as a group and compare that distribution across training. The source does not treat any one distribution shape as a universal diagnosis; the value is in seeing internal behavior that final loss alone cannot display.

Exploring Learned Embeddings

The Embeddings view focuses on locations and spatial relationships learned by an embedding layer. In the source example, the initial embedding layer learns representations for words in the input vocabulary. Because that embedding space has 128 dimensions, TensorBoard reduces it to 2D or 3D for inspection by using a selected dimensionality-reduction method: PCA or t-SNE.

nearbynearbyPositive word Aprojected pointPositive word Bprojected pointNegative word Aprojected pointNegative word Bprojected pointOutlier wordisolated projected point
How are high-dimensional learned embeddings arranged for inspection, and which examples form clusters or appear as outliers?

Interpreting the Embeddings View

A word embedding is projected from its original high-dimensional space into a 2D or 3D view. The displayed points include two visible groups and some points away from those groups. What can the developer inspect?

Inspect proximity: Look at which learned representations appear near one another in the displayed space.

Inspect clusters: Check whether examples form visible groups. In the source visualization, words with positive connotations and words with negative connotations form two visible clusters.

Inspect isolated points: Notice examples that do not appear near a visible group. Such points are worth examining as part of the representation analysis, without assuming a universal explanation from the picture alone.

Remember the training objective: Interpret the arrangement in the context of the task that trained the embedding. An embedding learned jointly with one objective is shaped by that objective and is not automatically a generic representation for another objective.

The Embeddings view reveals learned spatial relationships and task-shaped structure that a final loss value does not show.

Two Views of One Model

ViewWhat it representsMain use
TensorBoard Graphs tabLow-level TensorFlow operations underlying the Keras modelInspect the operational graph, including structures associated with gradient descent
Keras plot_modelConnected Keras layers used to define the modelSee a cleaner layer-level picture
can includecan displayTensorFlowoperationslow-level graphGradient descentstructurespart of operational graphKeras layerslayer-level graphShape informationshow_shapes option
How does the same model appear differently when represented as individual TensorFlow operations versus connected Keras layers?

TensorBoard's Graphs tab and Keras's plot_model utility are not duplicate displays. TensorBoard shows the low-level TensorFlow operations underlying a Keras model, and that graph can be much more complicated than the short layer sequence used to define the model. Keras's keras.utils.plot_model writes a cleaner layer-level graph; passing show_shapes=True adds shape information to that graph.

Mistakes in Interpretation

  • Treating the final metric as the complete story

    The metrics view is designed to show how measurements develop during a run, and internal views can reveal activation, gradient, or embedding behavior that the final value does not contain.

    Fix: Inspect metric trajectories and use histograms, embeddings, or graph views when the experiment requires more evidence.

  • Assuming a histogram is another kind of scalar metric

    A histogram shows a distribution of activation values or gradients, while a metric graph follows recorded measurements over the run.

    Fix: Use metric graphs for development over time and histograms for distributional internal behavior.

  • Confusing the Graphs tab with a layer diagram

    TensorBoard displays low-level TensorFlow operations and may include structures associated with gradient descent.

    Fix: Use keras.utils.plot_model for a cleaner layer-level picture, with show_shapes=True when shape information is useful.

  • Treating an embedding projection as a universal representation

    The source notes that embeddings learned jointly with a particular objective are shaped by that task.

    Fix: Interpret clusters and spatial relationships in the context of the objective that produced the embedding.

  • Assuming TensorBoard creates the training evidence by itself

    The training run produces log events, and TensorBoard reads those events to present visual views.

    Fix: Make the logging handoff explicit: training writes events, then TensorBoard reads them.

Practice the Inspection Cycle

MEDIUM

Imagine that a TensorFlow training run has written events and TensorBoard is open. Describe which TensorBoard view you would choose for each question: How did training and validation measurements develop over time? What are the distributions of layer activations? Where are learned word representations located relative to one another? What low-level operations underlie the model? Which view would you choose for a cleaner layer-level diagram?

Hints
  • Match changing measurements with the metrics view.
  • Match distributions of activations or gradients with histograms.
  • Match learned spatial relationships with the Embeddings view.
  • Separate TensorBoard's low-level operation graph from Keras's layer-level plot_model output.
  1. TensorBoard supports an iterative cycle: run an experiment, record events, inspect the resulting views, and use the observations to shape the next experiment. Metrics graphs show training and validation measurements over time. Histograms expose distributions of layer activations and gradients. Embedding visualizations show spatial relationships after reducing a high-dimensional space to 2D or 3D. TensorBoard's Graphs tab shows low-level TensorFlow operations, while Keras's plot_model provides a cleaner layer-level alternative.

Key Takeaways

  • TensorBoard turns information recorded during TensorFlow training into browser-based views for inspection.
  • Training and validation metric graphs show how recorded measurements develop during a run rather than only reporting a final value.
  • Histograms reveal distributions of layer activations and gradients, while embedding views reveal spatial relationships in learned representations.
  • TensorBoard's Graphs tab displays low-level TensorFlow operations; Keras's plot_model displays a cleaner layer-level graph and can include shape information.
  • Embedding patterns must be interpreted in the context of the objective that trained the embedding.