In this application, I will use SVD to find the principal components for a data set. Principal components means the axis (a vector) that represents the direction along which the data is most widely distributed, i.e the axis with the largest variance. In this example, I will use a 200 x 2 matrix which represents the 200 data points as plotted below.
I will find the same principal axes in three ways. The first applies SVD to the data matrix itself, the second applies SVD to DTD, and the third computes the eigenvectors of DTD. All three give the same axes, and comparing them shows how SVD and eigen decomposition are connected.
- Method 1 - SVD of the data matrix D
- Method 2 - SVD of DTD
- Method 3 - Eigen decomposition of DTD
- How do the three methods relate ?
- Matlab code for the three methods
- Reference
Method 1 - SVD of the data matrix D
Let's start with the most direct method. The 200 data points go into the rows of D, one point per row, and we take the SVD of D. The columns of V, which are the rows of transpose(V), are then the principal axes. No covariance matrix is needed.

First I assinged the D matrix to Matrix M. There is no mathmatical reason to do this. I just did this because I wanted to use the matrix M for all the matlab examples for this page.
![]()
Then I did SVD as follows.
![]()
D in the following plots represents the data set in scatter plot. You would intuitively find an axis along which the data are spread most widely. tr(V) shows two row vectors in transpose(V) vector that came from SVD. You see the red arrow is pointing the princial components, i.e the major axis. The third plot shows the transpose(V) vectors overlayed onto the data set. (Matlab source code for this example is in < List 1 >)
|
|
|||
The result of SVD is as follows. |
|||
|
M = D (data points) |
U |
S |
transpose(V) |
|
200 x 2 Matrix (Data Points) |
200 x 200 Matrix |
200 x 2 Matrix ------------------------ 14.6873 0 0 4.9745 . . all zeros . |
-0.4940 0.8695 0.8695 0.4940 |
The first row of tr(V) is the principal axis : (-0.4940, 0.8695) points at 119.6 degrees from the x axis. The listing generates the data with its long axis along y and then rotates it by 30 degrees, so the expected direction is 120 degrees.The singular values measure the spread : σ12 / (N-1) = 215.72 / 199 = 1.084 is the variance along the principal axis. σ22 / 199 = 0.124 is the variance across it.U is large but rarely needed : U is 200 x 200 here, and only its first two columns matter. D V = U S gives the coordinates of each point along the two axes, which are the principal component scores. The economy form svd(M,'econ') returns only those two columns.
Method 2 - SVD of DTD
The second method first squeezes the data into a 2 x 2 matrix. DTD has one row and one column per variable, whatever the number of data points. Its SVD must therefore give the same axes as Method 1, and the numbers below show that it does.
In this example, I defined the M matrix as follows. (The scatter matrix of the data set, DTD, is assinged to M. It equals N-1 = 199 times the covariance matrix when each column of D has zero mean. Now M is 2 x 2 square matrix)
![]()
Then I did SVD as follows.
![]()
D in the following plots represents the data set in scatter plot. You would intuitively find an axis along which the data are spread most widely. tr(V) shows two row vectors in transpose(V) vector that came from SVD. You see the red arrow is pointing the princial components, i.e the major axis. The third plot shows the transpose(V) vectors overlayed onto the data set. (Matlab source code for this example is in < List 2 >)
|
|
|||
Following is the result of SVD. You see that transpose(V) is same as the first example. S (eigen value) shows different value. It means the principal axis is same as in previous method. |
|||
|
M = tr(D).D |
U |
S |
transpose(V) |
|
71.3471 -82.0238 -82.0238 169.1167 |
-0.4940 0.8695 0.8695 0.4940 |
215.7181 0 0 24.7457 |
-0.4940 0.8695 0.8695 0.4940 |
The singular values are squared : 14.68732 = 215.72 and 4.97452 = 24.75. Because D = U S VT, DTD = V STS VT, so the new S holds the squares of the old singular values. The caption's S (eigen value) is right in this sense, because for DTD the singular values and the eigenvalues are the same numbers.U equals V here : DTD is symmetric positive semidefinite, so its SVD is also its eigen decomposition, and U = V. The U and transpose(V) columns of the table hold the same matrix, because this V is also symmetric.Scaling does not move the axes : dividing M by N-1 = 199 gives [0.3585 -0.4122 ; -0.4122 0.8498]. That matrix has the same eigenvectors as M, so the covariance matrix and DTD give the same principal axes.
Method 3 - Eigen decomposition of DTD
The third method skips SVD and asks for the eigenvectors of the same 2 x 2 matrix. For a symmetric positive semidefinite matrix, eigen decomposition and SVD must agree, so this is a check on Method 2. Only the order and the signs of the vectors differ.
In this example, I defined the M matrix as follows. (The scatter matrix of the data set, DTD, is assinged to M. It equals N-1 = 199 times the covariance matrix when each column of D has zero mean. Now M is 2 x 2 square matrix)
![]()
Then I calculated the eigenvector and eigen value.
D in the following plots represents the data set in scatter plot. You would intuitively find an axis along which the data are spread most widely. V shows two eigen vectors. Even though the color of the vector are swapped comparing to previous method, the direction of two arrows are still same. (Matlab source code for this example is in < List 3 >)
|
|
|||
|
Following is the result of Eigenvector and Eigenvalues. In case of SVD, the basis vector (vectors in V matrix) is arranged in decreasing order of the corresponding eigenvalues, but eigenvalue calculation does not follows this kind of ordering. |
|||
|
M = tr(D).D |
V(Eigenvector) |
E (Eigenvalues) |
|
|
71.3471 -82.0238 -82.0238 169.1167 |
-0.8695 -0.4940 -0.4940 0.8695 |
24.7457 0 0 215.7181 |
|
eig sorts in increasing order : Matlab returns 24.7457 first and 215.7181 second. So the principal axis is the second column of V, (-0.4940, 0.8695), which is the blue arrow in the plot above.An eigenvector has no fixed sign : the first column, (-0.8695, -0.4940), points opposite to the second row of tr(V) in Method 1. Both describe the same axis, because v and -v lie on the same line.The eigenvalues are the squared singular values of D : 215.7181 and 24.7457 match Method 2 exactly, and they match the squares of 14.6873 and 4.9745 from Method 1.
How do the three methods relate ?
The three tables give one pair of axes three times. Let's put the numbers side by side, and then see why they must agree and when the choice between them matters.
Method 1: svd(D) Method 2: svd(DTD) Method 3: eig(DTD) principal axis ( -0.4940, 0.8695 ) ( -0.4940, 0.8695 ) ( -0.4940, 0.8695 ) second axis ( 0.8695, 0.4940 ) ( 0.8695, 0.4940 ) ( -0.8695, -0.4940 ) value on axis 1 σ1 = 14.6873 215.7181 215.7181 value on axis 2 σ2 = 4.9745 24.7457 24.7457
Two details in the listings are worth knowing before you reuse them. The first is that rng(1) is called before rand and again before randn. So t and r are produced from the same random sequence, not from independent ones. Call rng(1) once if you want independent draws. The second is that the listings compute sigma1 and sigma2 but do not use them. The arrows therefore have unit length and show only direction. If you scale each arrow by sigma / sqrt(N-1), its length becomes the standard deviation along that axis, about 1.04 and 0.35.
The link is one line of algebra. If D = U S VT, then DTD = V STS VT. This is both an SVD and an eigen decomposition of DTD, with the same V and with eigenvalues σi2. So the three methods must return the same axes, up to order and sign.
The share of the variance on each axis follows from the same numbers. The principal axis carries 215.7181 / (215.7181 + 24.7457) = 89.7 % of the total. So a projection onto that single axis keeps about 90 % of the spread of the data. The Orthogonal Projection page shows that projection.
Two cautions apply in practice. First, PCA assumes that the data is centered. The listings never subtract the mean, and that is harmless only when the mean of each column is close to zero. For real data, subtract the column means from D first. Otherwise the first axis points toward the mean instead of along the spread. Second, Method 1 is the numerically preferred one. Forming DTD squares the condition number, here from 14.6873 / 4.9745 = 2.95 to 8.72, and for badly conditioned data this loses precision.
All three methods find the same axes : V from svd(D), V from svd(DTD) and the eigenvectors of DTD are the same vectors, up to order and sign.Singular values and eigenvalues are linked by a square : λi = σi2, and dividing by N-1 turns either one into a variance.Center the data first : without centering, the method describes the spread around the origin, not around the mean.Prefer the SVD of D : it avoids forming DTD, so it keeps the original condition number, and it also returns the scores U S directly.
Matlab code for the three methods
The three listings below produced the three figures above. They share the same data generator and differ only in the lines after M is defined. So you can switch between the methods by changing a few lines and compare the outputs.
< List 1 >
clear all;
N = 200;
rng(1);
t = rand(1,N);
rng(1);
r = randn(1,N);
x1 = 0.5 .* r .* cos(2*pi*t);
x2 =1.5 .* r .* sin(2*pi*t);
M = [x1; x2];
tm = [cos(pi/6) -sin(pi/6);sin(pi/6) cos(pi/6)]
D = (tm * M)';
x = D(:,1);
y = D(:,2);
M = D;
[U,S,V]=svd(M);
vT = V';
vT1 = vT(1,:);
vT2 = vT(2,:);
sigma1 = S(1,1);
sigma2 = S(2,2);
subplot(1,3,1);
plot(x,y,'ko','MarkerFaceColor',[0 0 0],'MarkerSize',2);
axis([-5 5 -5 5]);
title('D');
subplot(1,3,2);
quiver(0,0,vT1(1),vT1(2),'r','MaxHeadSize',1.5,'linewidth',2);
axis([-5 5 -5 5]);
hold on;
quiver(0,0,vT2(1),vT2(2),'b','MaxHeadSize',1.5,'linewidth',2);
axis([-5 5 -5 5]);
hold off;
title('tr(V)');
subplot(1,3,3);
plot(x,y,'ko','MarkerFaceColor',[0 0 0],'MarkerSize',2);
axis([-5 5 -5 5]);
hold on;
quiver(0,0,vT1(1),vT1(2),'r','MaxHeadSize',1.5,'linewidth',2);
axis([-5 5 -5 5]);
hold on;
quiver(0,0,vT2(1),vT2(2),'b','MaxHeadSize',1.5,'linewidth',2);
axis([-5 5 -5 5]);
hold off;
title('D and tr(V)');
< List 2 >
clear all;
N = 200;
rng(1);
t = rand(1,N);
rng(1);
r = randn(1,N);
x1 = 0.5 .* r .* cos(2*pi*t);
x2 =1.5 .* r .* sin(2*pi*t);
M = [x1; x2];
tm = [cos(pi/6) -sin(pi/6);sin(pi/6) cos(pi/6)]
D = (tm * M)';
x = D(:,1);
y = D(:,2);
M = D' * D;
[U,S,V]=svd(M);
vT = V';
vT1 = vT(1,:);
vT2 = vT(2,:);
sigma1 = S(1,1);
sigma2 = S(2,2);
subplot(1,3,1);
plot(x,y,'ko','MarkerFaceColor',[0 0 0],'MarkerSize',2);
axis([-5 5 -5 5]);
title('D');
subplot(1,3,2);
quiver(0,0,vT1(1),vT1(2),'r','MaxHeadSize',1.5,'linewidth',2);
axis([-5 5 -5 5]);
hold on;
quiver(0,0,vT2(1),vT2(2),'b','MaxHeadSize',1.5,'linewidth',2);
axis([-5 5 -5 5]);
hold off;
title('tr(V)');
subplot(1,3,3);
plot(x,y,'ko','MarkerFaceColor',[0 0 0],'MarkerSize',2);
axis([-5 5 -5 5]);
hold on;
quiver(0,0,vT1(1),vT1(2),'r','MaxHeadSize',1.5,'linewidth',2);
axis([-5 5 -5 5]);
hold on;
quiver(0,0,vT2(1),vT2(2),'b','MaxHeadSize',1.5,'linewidth',2);
axis([-5 5 -5 5]);
hold off;
title('D and tr(V)');
< List 3 >
clear all;
N = 200;
rng(1);
t = rand(1,N);
rng(1);
r = randn(1,N);
x1 = 0.5 .* r .* cos(2*pi*t);
x2 =1.5 .* r .* sin(2*pi*t);
M = [x1; x2];
tm = [cos(pi/6) -sin(pi/6);sin(pi/6) cos(pi/6)]
D = (tm * M)';
x = D(:,1);
y = D(:,2);
M = D' * D;
[V,E] = eig(M);
vT = V;
vT1 = vT(:,1);
vT2 = vT(:,2);
sigma1 = E(1,1);
sigma2 = E(2,2);
subplot(1,3,1);
plot(x,y,'ko','MarkerFaceColor',[0 0 0],'MarkerSize',2);
axis([-5 5 -5 5]);
title('D');
subplot(1,3,2);
quiver(0,0,vT1(1),vT1(2),'r','MaxHeadSize',1.5,'linewidth',2);
axis([-5 5 -5 5]);
hold on;
quiver(0,0,vT2(1),vT2(2),'b','MaxHeadSize',1.5,'linewidth',2);
axis([-5 5 -5 5]);
hold off;
title('V');
subplot(1,3,3);
plot(x,y,'ko','MarkerFaceColor',[0 0 0],'MarkerSize',2);
axis([-5 5 -5 5]);
hold on;
quiver(0,0,vT1(1),vT1(2),'r','MaxHeadSize',1.5,'linewidth',2);
axis([-5 5 -5 5]);
hold on;
quiver(0,0,vT2(1),vT2(2),'b','MaxHeadSize',1.5,'linewidth',2);
axis([-5 5 -5 5]);
hold off;
title('D and V');
Reference
[1] Principal Component Analysis (PCA) - YouTube
[2] StatQuest: Principal Component Analysis (PCA), Step-by-Step - YouTube


