Linear Transformations
Dimensionality reduction maps vectors from R^d to R^n using a linear transformation, where n is smaller than d.
Why Fewer Coordinates Can Help
Suppose each vector in a dataset has d features. Working with every feature can be unnecessary when a smaller representation can retain the important structure of the data. Dimensionality reduction addresses this by replacing each vector with a representation that has fewer dimensions. In PCA, this replacement is a linear transformation followed by a reconstruction process, so the reduced representation is judged by how closely the original vectors can be approximated.
Dimensionality reduction maps vectors from R^d to R^n using a linear transformation, where n is smaller than d.
The Compression Step
Start with one vector x in R^d. The compression matrix W has n rows and d columns, with n smaller than d. Multiplying W by x produces y = W x, a vector in R^n. Because n is smaller than d, y contains fewer coordinates than x. This move from R^d to R^n is the central dimensionality-reduction step.
Tracking the dimensions
Consider an original vector x in R^4. Let W have 2 rows and 4 columns. Determine the space containing the compressed representation y = W x.
Identify the original vector: The vector x has four coordinates, so it belongs to R^4.
Inspect the compression matrix: W has 2 rows and 4 columns. Its 4 columns match the four coordinates of x, and its 2 rows determine the number of coordinates in the result.
Determine the result: Multiplying W by x produces y with two coordinates. Therefore y belongs to R^2.
The transformation compresses x from R^4 into y in R^2. The representation has fewer coordinates than the original.
The Recovery Step
Compression alone gives the smaller vector y. PCA also describes how to return to the original dimensional setting. The recovery matrix U has d rows and n columns. Multiplying U by y produces the recovered vector, written as x̃ = U y. This recovered vector lies in R^d, just as the original x does, but it is generally an approximation rather than a claim that every detail of x has been restored exactly.
Returning to the original space
A compressed vector y belongs to R^2. Let U have 4 rows and 2 columns. Determine the space containing the recovered vector x̃ = U y.
Identify the compressed input: The compressed representation y has two coordinates and belongs to R^2.
Inspect the recovery matrix: U has 2 columns, matching the two coordinates of y. It has 4 rows, so its output has four coordinates.
Determine the recovered space: The product U y is a four-coordinate vector in R^4.
Interpret the result: The recovered vector has the original dimensionality, but it is an approximation because the compression used fewer dimensions.
U maps the compact representation from R^2 back into R^4, producing an approximation of the original vector.
When tracing a PCA transformation, track both the space and the role of each object: x is the original vector, W performs compression, y is the compact representation, U performs recovery, and x̃ is the recovered approximation.
What PCA Optimizes
The matrices W and U are not selected merely because their shapes are compatible. PCA looks for the pair that makes recovered vectors close to the original vectors across the dataset. If the dataset contains vectors x1 through xm, each vector is compressed and recovered as x̃i = U W xi. PCA chooses W and U to minimize the total squared distance between each original vector xi and its recovered version x̃i.
PCA's reconstruction objective is to minimize the total squared distance between the original dataset vectors and the vectors recovered after compression and recovery.
The objective connects both transformations. W controls the move into the smaller space, U maps the compact representation back into the original space, and the squared distance evaluates the combined result. A pair of matrices is useful in PCA only when this combined process keeps the recovered vectors close to the originals across the dataset.
Common Reasoning Mistakes
Treating compression as if it preserves the original number of coordinates.
The compression matrix W has n rows with n smaller than d, so y = W x belongs to R^n.
Fix:
Track the output dimension from the number of rows of W.Assuming recovery exactly restores the original vector.
The source describes x̃ as an approximation and evaluates it by its distance from x.
Fix:
Use the phrase recovered approximation and compare it with the original using squared distance.Describing PCA as compression without an accuracy objective.
PCA chooses W and U to minimize the total squared distance between original and recovered vectors.
Fix:
Explain both stages: dimensionality reduction and reconstruction-error minimization.Checking only whether matrix dimensions are compatible.
Compatible shapes allow the operations, but PCA selects the matrices according to how closely the recovered vectors match the originals.
Fix:
Separate shape compatibility from the optimization goal.
Trace the Transformation
A vector x belongs to R^5. A compression matrix W has 2 rows and 5 columns. A recovery matrix U has 5 rows and 2 columns. Identify the space of y = W x, the space of x̃ = U y, and the quantity PCA tries to minimize across a dataset.
Hints
- For W x, use the number of rows of W to determine the output dimension.
- For U y, use the number of rows of U to determine the output dimension.
- The objective compares every original vector with its recovered approximation.
- Dimensionality reduction maps an original vector from R^d into a smaller space R^n, where n is less than d. The compression matrix W has n rows and d columns and produces y = W x. The recovery matrix U has d rows and n columns and produces the approximation x̃ = U y in R^d. PCA chooses W and U by minimizing the total squared distance between original vectors and their recovered approximations.
Key Takeaways
- Dimensionality reduction is a linear mapping from R^d to R^n, with n smaller than d.
- The compression matrix W produces the compact representation y = W x.
- The recovery matrix U maps y back into R^d as an approximation x̃ = U y.
- PCA selects W and U to minimize the total squared distance between original and recovered dataset vectors.
- A recovered vector can have the original number of coordinates without restoring every original detail.