Concepts / Image Editing with Concept Vectors

Image Editing with Concept Vectors

A VAE combines an encoder, a sampling step, and a decoder.

  • Programming

From Image to Reconstruction

A variational autoencoder, or VAE, transforms an input image through three connected stages. The encoder describes the image with parameters for a latent distribution. The sampling step uses those parameters and random noise to choose a latent point. The decoder then turns that sampled point into a reconstructed image. The important difference from a system with one fixed middle code is that the VAE represents uncertainty and variation in latent space rather than recording only one point.

encodesample from distributiondecode zInput imageimage dataEncoderdistribution parametersSampling steplatent point zDecoderreconstructed image
How does an input image move through the encoder, sampling step, and decoder to become a reconstructed image?

The Encoder's Two Outputs

The encoder does not produce two images or two alternative reconstructions. It maps the input image to two vectors that describe a distribution in latent space. In the implementation described by the source, convolutional layers process the image, the result is flattened, and a dense layer creates an intermediate representation. Two further dense layers produce z_mean and z_log_var. The source also refers to the second quantity as z_log_variance. Together, these outputs describe where the latent distribution is centered and how its spread is represented.

encodeencodeencodecombinecombineInput imageone fixed codeInput imagedistribution descriptionFixed latent codeone pointz_meancenterz_log_variancespread informationSampled point zchosen latent point
What is the difference between encoding an image as one fixed point and encoding it as a probability distribution in latent space?

The decoder receives the sampled point z. It does not decode the original image directly, and it does not decode z_mean or z_log_variance by themselves.

Sampling a Latent Point

The sampling stage combines the encoder's distribution information with randomness. z_mean supplies the encoded location, z_log_variance supplies variance-related information, and epsilon is a random tensor of small values. Their combination produces z, the latent-space point passed to the decoder. This means that repeated samples can be points around the encoded location rather than one permanently fixed middle representation.

locationspread informationrandom variationdecodez_meanencoded locationz_log_variancevariance-related parameterepsilonrandom small valueszsampled latent pointDecoderreconstructed image
How are z_mean and z_log_variance combined with random noise epsilon to produce the sampled latent vector z?

Tracing One Image Through Sampling

Follow the roles of the four named quantities after one image reaches the encoder.

1. Encode: The encoder processes the image and produces z_mean together with z_log_variance. These are distribution parameters, not reconstructed images.

2. Add random input: The sampling operation receives epsilon, a random tensor of small values, along with the two encoder outputs.

3. Form the latent point: The operation combines z_mean, z_log_variance, and epsilon to produce z.

4. Decode: The decoder receives z and uses it to produce the reconstructed image.

The path is input image to z_mean and z_log_variance, then through epsilon-based sampling to z, then from z to the reconstructed image.

Training Two Useful Pressures

A VAE is trained with two loss functions that provide different signals. Reconstruction loss encourages the decoded sample to match the original input image. Regularization loss encourages a well-formed latent space and helps reduce overfitting to the training data. The two pressures work together: reconstruction helps the model reproduce inputs, while regularization supports the organized latent structure that makes sampling and manipulation useful.

comparecompareorganizetraining signaltraining signalOriginal inputtraining imageDecoded samplereconstructed imageReconstruction lossmatch the inputLatent spaceorganized structureRegularization losswell-formed spaceVAE trainingcombined pressure
How do reconstruction loss and regularization loss provide different training signals, and how do they work together?
LossMain questionTraining role
Reconstruction lossDoes the decoded sample match the original input?Encourages reproduction of the input
Regularization lossIs the latent space well formed?Supports organized structure and helps reduce overfitting

Moving Through Latent Space

A continuous and structured latent space makes image editing possible through concept vectors. A concept vector represents a direction in latent space associated with changing an aspect of the data. Moving a latent point along such a direction and decoding the new point can change the corresponding aspect of the decoded image. The source supports the idea that directions can become meaningful for changing aspects of the data; it does not imply that every movement changes only one attribute or that other visual properties are guaranteed to remain unchanged.

decodemove alongguide movementdecodezoriginal latent pointDecoded imageoriginal decoded resultConcept vectorlatent directionMoved latent pointnew pointEdited decoded imageresult after decoding
How does moving along a concept direction in latent space affect the decoded image?

Imagine that a trained latent space contains a direction associated with one visual aspect of its data. Starting with z, an editing process can move to a nearby latent point in that direction and send the new point through the decoder. The decoded result may show the associated change because nearby points are encouraged to produce similar images and meaningful directions can support manipulation.

Mistakes in the VAE Trace

  • Treating z_mean and z_log_variance as two images.

    These outputs are vectors describing a latent distribution. They are not images or alternative reconstructions.

    Fix: Treat them as distribution parameters that participate in producing the sampled latent point z.

  • Sending the original image or distribution parameters directly to the decoder.

    The sampled point z is the quantity passed to the decoder.

    Fix: Trace the middle stage explicitly: encoder outputs, sampling with epsilon, then z, then decoding.

  • Adding randomness without explaining its purpose.

    The random sampling process helps organize the latent space so nearby points can decode into similar images.

    Fix: Connect randomness to continuity and to the usefulness of latent-space manipulation.

  • Using reconstruction loss as the whole training objective.

    Reconstruction loss focuses on reproducing inputs, while regularization supports a well-formed latent space and helps reduce overfitting.

    Fix: Remember that both reconstruction and regularization pressures matter.

Check Your Understanding

MEDIUM

A VAE has processed an image. Explain the job of each item in this sequence: z_mean, z_log_variance, epsilon, z, and the decoder. Then explain what could be lost if training used reconstruction loss but no regularization loss.

Hints
  • Start with the two outputs produced by the encoder.
  • Identify which quantity is random and which quantity is passed to the decoder.
  • Relate reconstruction loss to matching the input and regularization loss to latent-space organization.

What do you think happens?

Which item is decoded by the VAE: z_mean, z_log_variance, or z?

  • z_mean
  • z_log_variance
  • z
Reveal answer

Answer: z

The encoder produces z_mean and z_log_variance, the sampling step combines them with epsilon to produce z, and z is the latent-space point passed to the decoder.

Key Takeaways

  1. A VAE has three stages: encoding, sampling, and decoding.
  2. The encoder produces z_mean and z_log_variance to describe a latent distribution rather than one fixed code.
  3. epsilon supplies random variation, and the sampling step uses it with the encoder outputs to produce z.
  4. The decoder reconstructs the image from z, not directly from the original image or the distribution parameters.
  5. Reconstruction loss supports input matching, while regularization loss supports an organized latent space useful for sampling and image editing.

Key Takeaways

  • A VAE transforms an image through encoding, sampling, and decoding.
  • Its encoder describes a latent distribution with z_mean and z_log_variance instead of producing only one fixed latent code.
  • epsilon helps create the sampled latent point z, which is the value decoded into an image.
  • Reconstruction loss and regularization loss provide complementary training pressures.
  • A continuous, structured latent space makes concept-vector movement useful for image editing.