A covariance matrix packs the spread of several data columns, and the way those columns move together, into one square matrix. It is the starting point for principal component analysis, for MMSE estimation and for most MIMO channel models. I'll build it up from the variance of a single column, then list its properties, and then use scatter plots to show what its eigenvectors and eigenvalues mean. The last sections compare covariance with correlation and explain why the matrix form is so useful.
- How are variance and covariance defined for data columns ?
- Properties of Covariance Matrix
- Graphical Implication of Covariance Matrix
- Covariation vs. Correlation
- Why Covariance Matrix ?
How are variance and covariance defined for data columns ?
Before going into the mathematical definition for general case, let's start with specific cases first. The variance measures how far one column of data spreads around its own mean. The covariance measures whether two columns move up and down together.
Let's assum that we have a single column data as shown below.

Variance of the data X can be defined as follows : (As you see, in case of single column data, Variance is same as covariance between the same data).

Now let's assume that we have two set of single column data as shown below.

If we calculate the variance of each data set separately, it would be as follows.

The covariance between these two data set (two single column data) is defined as follows :

If you combine the two sets of data into a matrix form as shown below

The covariance comes out as a matrix as shown below. (Since it has 2 columns of data, the covariance matrix becomes 2 x 2 matrix)

If the average of each data set (each column) is zero, the covariance matrix of the matrix can be calculated as follows. (AWGN (Additive White Gaussian Noise) is a good example of this kind of data)

Two details decide whether your numbers match the output of a tool. The first is the divisor. The formulas above divide by N, which gives the population covariance. Matlab cov(), numpy cov() and most statistics packages divide by N-1 by default, which gives the sample covariance. For a few hundred samples the difference is below one percent, but for a handful of samples it is large. The second detail is complex data. When the columns of A are complex, as with I/Q samples, the formula becomes AHA/N. It uses the conjugate transpose instead of the plain transpose.
Let's check the definitions on a small data set with N = 4. Take X = [2, 4, 6, 8]T and Y = [1, 3, 2, 6]T. The means are 5 and 3, so the deviations are [-3, -1, 1, 3] and [-2, 0, -1, 3]. The sums of products then give the three entries below.
Var(X) = (9 + 1 + 1 + 9) / 4 = 5
Var(Y) = (4 + 0 + 1 + 9) / 4 = 3.5
Cov(X,Y) = (6 + 0 - 1 + 9) / 4 = 3.5
Cov(A) = [ 5 3.5 ] divide by N = 4
[ 3.5 3.5 ]
cov(X,Y) = [ 6.6667 4.6667 ] Matlab / numpy default, divide by N-1 = 3
[ 4.6667 4.6667 ]
Variance is the covariance of a column with itself : that is why the diagonal of Cov(A) holds the variances.The sign of the covariance gives the direction of the relation : a positive value means the two columns tend to rise together. A negative value means one column tends to fall when the other rises.Check the divisor before you compare numbers : the formulas on this page divide by N, while cov() in Matlab and numpy divides by N-1.ATA/N needs zero-mean columns : subtract the column means first. Otherwise the result mixes the means into the covariance.
Properties of Covariance Matrix
The definitions above already fix the shape and the symmetry of the matrix. The list below collects the basic properties, and the paragraph after it adds the ones that matter when you analyze the eigenvalues and eigenvectors of the matrix.
- Covariance matrix can represents the variance and linear correlation in multivariate/multidimensional data
- Covariance matrix gives you meaningful result only when the data sets have linear correlation.
- Covariance matrix is always a Square Matrix
- If Idata is n rows and m columns (n x m Matrix), the dimension of Covariance matrix is m x m Square matrix
Three more properties follow directly from the definition. The matrix is symmetric, because Cov(X,Y) and Cov(Y,X) are the same sum. It is also positive semidefinite, so none of its eigenvalues is negative. The reason is that wTCov(A)w is the variance of the combined column Aw, and a variance can never be negative. Finally, its trace is the total variance, the sum of the variances of all columns, and the trace equals the sum of the eigenvalues. In the worked example above, the trace is 5 + 3.5 = 8.5, and the eigenvalues are about 7.83 and 0.67.
The matrix is square and symmetric : for m data columns it is m x m, and entry (i, j) equals entry (j, i). For complex data it is Hermitian instead.No eigenvalue is negative : each eigenvalue is the variance of the data along its eigenvector.The eigenvectors are orthogonal : because the matrix is symmetric, they form a set of perpendicular axes for the data. The next section shows those axes on scatter plots.
Graphical Implication of Covariance Matrix
A scatter plot shows what the covariance matrix encodes. The eigenvectors of the matrix point along the main axes of the point cloud, and each eigenvalue gives the variance of the data along its eigenvector. The Matlab code below produces a cloud of this kind. It draws 200 points with standard deviations 0.25 and 1. Then it rotates the cloud so that its long axis lies along the 45 degree diagonal, and it computes the covariance matrix and its eigen decomposition.
X = 0.25*randn(1,200); Y = randn(1,200); A = [cos(pi/4) sin(pi/4);-sin(pi/4) cos(pi/4)]; Tx = A * [X;Y]; X = Tx(1,:); Y = Tx(2,:); c=cov(X,Y) plot(X,Y,'ro','MarkerFaceColor',[1 0 0],'MarkerSize',2.0); axis([-5 5 -5 5]); [v,l]=eigs(c) v * l * inv(v)
The last line, v * l * inv(v), rebuilds the covariance matrix from its eigenvectors and eigenvalues. That is the decomposition shown at the end of each example. The three examples below use three different point clouds. Each one shows the scatter plot with the eigenvectors, then the covariance matrix, the eigenvectors and the eigenvalues. In all three examples, λ1 is the larger eigenvalue.
Example 1
In this cloud the points spread about equally in every direction, with no visible tilt. The covariance matrix is therefore close to the identity matrix, with variances near 1 and a small covariance of -0.0978.






The eigenvalue listing in this example shows λ2 = 0.0904, but that value is a typo. The eigenvalues of the covariance matrix shown are 1.1053 and 0.9094. The trace confirms it: 1.1053 + 0.9094 = 2.0147, which is the sum 1.0020 + 1.0127 of the diagonal entries. With two nearly equal eigenvalues, the directions of the eigenvectors carry little meaning. A small change in the random data can turn them to almost any angle.
Example 2
Here the points spread mainly along the vertical axis. The variance of Y, 0.9444, is much larger than the variance of X, 0.0601. The covariance is small, so the matrix is close to diagonal. The two variances are close to 0.0625 and 1, the values that the first two lines of the code produce before the rotation.





The eigenvector v1 = [0.0302, 0.9995]T is almost the vertical axis, and its eigenvalue 0.9452 is almost the variance of Y. The second eigenvector is almost the horizontal axis, with eigenvalue 0.0593. So when the covariance is small compared with the variances, the eigenvectors stay close to the coordinate axes.
Example 3
This cloud has the shape that the code above produces, with its long axis along the 45 degree diagonal. The two variances are nearly equal, and the covariance 0.5071 is almost as large as either of them.






The eigenvector v1 = [0.7058, 0.7084]T points along the diagonal. Its eigenvalue 1.0681 is close to 1, the variance of Y before the rotation, and the eigenvalue 0.0538 is close to 0.0625, the variance of 0.25*randn. So the eigen decomposition undoes the rotation. It recovers both the axes and the variances of the original, unrotated data, and this is the idea behind principal component analysis.
Eigenvectors give the axes of the cloud : v1 points along the direction of the largest spread.Eigenvalues give the variance along each axis : a large ratio between them means a long, thin cloud, and nearly equal values mean a round one.The trace checks the eigenvalues : the eigenvalues always add up to the sum of the diagonal entries.
Covariation vs. Correlation
Covariance has the units of X times the units of Y, so its size depends on the scale of the data. If you measure X in millivolts instead of volts, the covariance grows by a factor of 1000. The correlation removes that dependence. It divides the covariance by both standard deviations, as the formula below shows.

The correlation always lies between -1 and +1. A value of +1 or -1 means that the points lie exactly on a straight line, and 0 means no linear relation. The divisor N or N-1 appears in the covariance and in both standard deviations, so it cancels. The correlation is therefore the same for both conventions. For Example 3, the correlation is 0.5071 / √(0.5592 x 0.5628), which is about 0.90. For Example 1, it is about -0.10. So the correlation confirms what the plots show: a strong linear relation in Example 3 and almost none in Example 1. In Matlab, corrcoef(X,Y) returns the correlation matrix, and numpy provides np.corrcoef.
Covariance depends on scale, and correlation does not : compare correlations when the columns have different units or ranges.Correlation measures only a linear relation : if X is symmetric around zero and Y = X2, the correlation is 0, even though Y depends completely on X.The correlation matrix has ones on its diagonal : each column is perfectly correlated with itself.
Why Covariance Matrix ?
In most cases, Covariance matrix is applied to a long sequence of data set (i.e, multiple huge vectors) but it is not easy to get any useful information directly from the original data itself (we don't have much mathematical tools for it). But taking the covariance matrix from those dataset, we can get a lot of useful information with various mathematical tools that are already developed. This is possible mainly because of the following properties of covariance matrix.
- The covariance matrix is always square matrix (i.e, n x n matrix). We have many tools to extract useful informations from square matrix, like eigen values and eigen vectors.
- If the dataset is real valued data, the covariance matrix is in real symetric form. In this case, the eigen vector gives the orthogonal basis for those dataset (similar to SVD). See the example 1,2,3 in this page.
- If the dataset is complex numbered data, the covariance matrix is a Hermitian Matrix.
In wireless communication, the same matrix appears under several names. The noise covariance E[nnH] = σ2I describes AWGN on several receive antennas. Its diagonal form says that the antennas see independent noise of equal power. The channel covariance E[hhH] describes how strongly the antennas are correlated, and it is an input to MMSE channel estimation. The MMSE equalizer (HHH + σ2I)-1HH also carries the noise covariance σ2I next to HHH. And principal component analysis applies the eigen decomposition of Example 3 to data with many columns.
The matrix turns a long data record into a small square matrix : 200 samples of two columns become four numbers, and the tools for square matrices then apply.Its eigen decomposition is always well behaved : the eigenvalues are real and never negative, and the eigenvectors are orthogonal.Complex data needs the Hermitian form : use AHA/N for I/Q samples, so that the diagonal still holds real, non-negative powers.