Autoencoders
A VAE combines an encoder, a sampling step, and a decoder.
From Images to Latent Points
A variational autoencoder, or VAE, learns a compact space of image representations. Instead of sending an image through an encoder into one permanent code, it describes a distribution of possible latent points, samples one point from that distribution, and sends the sampled point through a decoder. The decoder then produces an image.
A latent space of images is a low-dimensional vector space whose points can be mapped to realistic-looking images.
Encoding a Distribution
The encoder maps an input image to two vectors: z_mean and z_log_variance. These outputs describe a latent distribution. They are not two images, and they are not two alternative reconstructions. They are distribution-related parameters that determine how the model will sample a latent point.
| Encoder output | Role in the VAE |
|---|---|
| z_mean | Describes the central location of the latent distribution |
| z_log_variance | Provides a variance-related description used during sampling |
| z | The sampled latent point that is passed to the decoder |
Sampling the Latent Vector
The sampling stage combines z_mean and z_log_variance with epsilon, a random tensor of small values. This produces z, the latent-space point that the decoder receives. The original image, z_mean, and z_log_variance are not the values sent directly to the decoder in place of z.
Tracing One Image Through Sampling
Follow the information flow for one input image in a VAE.
Encode: The encoder processes the image and produces z_mean and z_log_variance.
Sample: The sampling step combines those distribution parameters with epsilon, a random tensor of small values, to produce z.
Decode: The decoder receives z and maps that latent point into an image grid of pixels.
The decoder uses the sampled latent point z, not the distribution parameters alone, to produce the reconstructed image.
Training Pressures
A VAE is trained with two loss functions that apply different pressures. Reconstruction loss encourages the decoded sample to match the original input. Regularization loss encourages a well-formed latent space and helps reduce overfitting to the training data.
| If this pressure dominated | Main focus |
|---|---|
| Reconstruction loss | Reproducing the training input |
| Regularization loss | Maintaining an organized latent space and reducing overfitting |
Decoding New Images
The decoder maps a selected or randomly sampled latent point into an image grid of pixels. During reconstruction, the point comes from sampling an encoded image. After the latent space has been learned, a point can also be selected deliberately or sampled at random and passed through the decoder.
Imagine that training has developed a latent space of image representations. Selecting a point in that space and sending it through the decoder produces an image in image space. Because the point can be selected rather than copied from one training image, the resulting image can be one the system has never seen before.
When explaining a VAE, always distinguish image space from latent space. The encoder moves from an image to a latent distribution, sampling selects a latent point, and the decoder moves from that point back to an image.
VAEs and GANs
The source material characterizes a VAE as learning a continuous, structured latent space in which nearby points can decode into similar images and directions can represent meaningful axes of variation. It identifies generative adversarial networks as a related concept, but it does not provide enough information here to make a detailed factual claim about how GAN latent spaces are trained or navigated.
Common Misunderstandings
Treating z_mean and z_log_variance as two images
These outputs describe the latent distribution used for sampling. They are not images or alternative reconstructions.
Fix:
Think of them as distribution parameters. The sampled point z is what the decoder receives.Sending the original image directly to the decoder
The VAE's middle stage is the sampled latent point, not the original image.
Fix:
Trace the path as image, distribution parameters, sampled z, then decoded image.Assuming randomness makes the latent space disorganized
The source explains that sampling around an encoded location encourages nearby latent points to produce similar images.
Fix:
Understand randomness as part of the mechanism that supports continuity.Focusing only on reconstruction loss
Reconstruction loss supports reproduction, but regularization supports a well-formed latent space and helps reduce overfitting.
Fix:
Consider both reconstruction and regularization pressures.
Check Your Understanding
A VAE receives an image. Explain, in order, what happens to the image, z_mean, z_log_variance, epsilon, z, and the decoder. Then explain what would be lost if training used reconstruction loss without regularization loss.
Hints
- Start with the encoder's two distribution-related outputs.
- Identify which value is actually passed to the decoder.
- Relate reconstruction loss to matching the input and regularization loss to organizing the latent space.
What do you think happens?
After a VAE has learned its latent space, what can happen if a new latent point is selected and passed through the decoder?
Reveal answer
Answer: The decoder can produce an image that the system has never seen before.
The latent space contains image representations, and the decoder maps a selected or randomly sampled latent point into an image grid of pixels.
Key Takeaways
- A VAE has three main stages: encoding, sampling, and decoding.
- The encoder produces z_mean and z_log_variance, which describe a latent distribution rather than one fixed code.
- The sampling step combines those parameters with epsilon to produce z, the latent point passed to the decoder.
- Reconstruction loss supports matching the input, while regularization loss supports an organized latent space.
- A decoder can map a deliberately selected or randomly sampled latent point to an image that was not present in the training data.
Key Takeaways
- A variational autoencoder transforms an image through encoding, sampling, and decoding.
- Its encoder describes a latent distribution with z_mean and z_log_variance instead of producing only one fixed latent code.
- epsilon contributes randomness during sampling, producing z, which is the point decoded into an image.
- Reconstruction and regularization losses work together to support both accurate reproduction and an organized latent space.
- Because a decoder maps latent points to images, selecting a new point can produce an image the model has not previously seen.