Concepts / Autoencoders

Autoencoders

A VAE combines an encoder, a sampling step, and a decoder.

  • Programming

From Images to Latent Points

A variational autoencoder, or VAE, learns a compact space of image representations. Instead of sending an image through an encoder into one permanent code, it describes a distribution of possible latent points, samples one point from that distribution, and sends the sampled point through a decoder. The decoder then produces an image.

A latent space of images is a low-dimensional vector space whose points can be mapped to realistic-looking images.

encodedescribe distributionpass zdecodeInput imageimage pixelsEncoderdistribution parametersSamplingrandom latent pointDecoderimage-producing networkReconstructed imageimage pixels
How does an input image move through the encoder, sampling step, and decoder to become a reconstructed image?

Encoding a Distribution

The encoder maps an input image to two vectors: z_mean and z_log_variance. These outputs describe a latent distribution. They are not two images, and they are not two alternative reconstructions. They are distribution-related parameters that determine how the model will sample a latent point.

Encoder outputRole in the VAE
z_meanDescribes the central location of the latent distribution
z_log_varianceProvides a variance-related description used during sampling
zThe sampled latent point that is passed to the decoder
encodeencodeencodesamplesampleInput imageimageInput imageimageOne codesingle latent vectorz_meandistribution locationz_log_variancevariance-related parameterzsampled latent point
What is the difference between an encoder producing one fixed latent vector and producing parameters for a latent distribution?

Sampling the Latent Vector

The sampling stage combines z_mean and z_log_variance with epsilon, a random tensor of small values. This produces z, the latent-space point that the decoder receives. The original image, z_mean, and z_log_variance are not the values sent directly to the decoder in place of z.

combinecombineadd randomnessz_meandistribution locationz_log_variancevariance-related parameterepsilonrandom tensorzsampled latent point
How are z_mean and z_log_variance used with random noise epsilon to calculate the sampled latent vector z?

Tracing One Image Through Sampling

Follow the information flow for one input image in a VAE.

Encode: The encoder processes the image and produces z_mean and z_log_variance.

Sample: The sampling step combines those distribution parameters with epsilon, a random tensor of small values, to produce z.

Decode: The decoder receives z and maps that latent point into an image grid of pixels.

The decoder uses the sampled latent point z, not the distribution parameters alone, to produce the reconstructed image.

Training Pressures

A VAE is trained with two loss functions that apply different pressures. Reconstruction loss encourages the decoded sample to match the original input. Regularization loss encourages a well-formed latent space and helps reduce overfitting to the training data.

pressurepressuresupportssupportsReconstruction lossmatch the original inputVAE trainingtwo pressuresInput reproductiondecoded imageRegularization losswell-formed latent spaceOrganized latentspacesampling and manipulation
How do reconstruction loss and regularization loss combine to make the decoder reproduce the input while keeping the latent space organized?
If this pressure dominatedMain focus
Reconstruction lossReproducing the training input
Regularization lossMaintaining an organized latent space and reducing overfitting

Decoding New Images

The decoder maps a selected or randomly sampled latent point into an image grid of pixels. During reconstruction, the point comes from sampling an encoded image. After the latent space has been learned, a point can also be selected deliberately or sampled at random and passed through the decoder.

place pointpass pointproduceSelect latent pointdeliberate or randomLatent spaceimage representationsDecodermaps point to pixelsPreviously unseenimageimage grid
How can choosing a new point in the latent space and passing it through the decoder produce an image that was not in the training set?

Imagine that training has developed a latent space of image representations. Selecting a point in that space and sending it through the decoder produces an image in image space. Because the point can be selected rather than copied from one training image, the resulting image can be one the system has never seen before.

When explaining a VAE, always distinguish image space from latent space. The encoder moves from an image to a latent distribution, sampling selects a latent point, and the decoder moves from that point back to an image.

VAEs and GANs

The source material characterizes a VAE as learning a continuous, structured latent space in which nearby points can decode into similar images and directions can represent meaningful axes of variation. It identifies generative adversarial networks as a related concept, but it does not provide enough information here to make a detailed factual claim about how GAN latent spaces are trained or navigated.

encouragessource detailVAEcontinuous structured spaceNearby pointssimilar decoded imagesGANrelated conceptNot specifiedin this source pack
How does the source material characterize the latent-space structure of a VAE compared with the information provided about GANs?

Common Misunderstandings

  • Treating z_mean and z_log_variance as two images

    These outputs describe the latent distribution used for sampling. They are not images or alternative reconstructions.

    Fix: Think of them as distribution parameters. The sampled point z is what the decoder receives.

  • Sending the original image directly to the decoder

    The VAE's middle stage is the sampled latent point, not the original image.

    Fix: Trace the path as image, distribution parameters, sampled z, then decoded image.

  • Assuming randomness makes the latent space disorganized

    The source explains that sampling around an encoded location encourages nearby latent points to produce similar images.

    Fix: Understand randomness as part of the mechanism that supports continuity.

  • Focusing only on reconstruction loss

    Reconstruction loss supports reproduction, but regularization supports a well-formed latent space and helps reduce overfitting.

    Fix: Consider both reconstruction and regularization pressures.

Check Your Understanding

MEDIUM

A VAE receives an image. Explain, in order, what happens to the image, z_mean, z_log_variance, epsilon, z, and the decoder. Then explain what would be lost if training used reconstruction loss without regularization loss.

Hints
  • Start with the encoder's two distribution-related outputs.
  • Identify which value is actually passed to the decoder.
  • Relate reconstruction loss to matching the input and regularization loss to organizing the latent space.

What do you think happens?

After a VAE has learned its latent space, what can happen if a new latent point is selected and passed through the decoder?

  • Only the exact training image used to create the point can be returned
  • A previously unseen image can be produced
  • The latent point must be sent back through the encoder first
Reveal answer

Answer: The decoder can produce an image that the system has never seen before.

The latent space contains image representations, and the decoder maps a selected or randomly sampled latent point into an image grid of pixels.

Key Takeaways

  1. A VAE has three main stages: encoding, sampling, and decoding.
  2. The encoder produces z_mean and z_log_variance, which describe a latent distribution rather than one fixed code.
  3. The sampling step combines those parameters with epsilon to produce z, the latent point passed to the decoder.
  4. Reconstruction loss supports matching the input, while regularization loss supports an organized latent space.
  5. A decoder can map a deliberately selected or randomly sampled latent point to an image that was not present in the training data.

Key Takeaways

  • A variational autoencoder transforms an image through encoding, sampling, and decoding.
  • Its encoder describes a latent distribution with z_mean and z_log_variance instead of producing only one fixed latent code.
  • epsilon contributes randomness during sampling, producing z, which is the point decoded into an image.
  • Reconstruction and regularization losses work together to support both accurate reproduction and an organized latent space.
  • Because a decoder maps latent points to images, selecting a new point can produce an image the model has not previously seen.