Engineering Math - Matrix

 

 

 

Orthonormality

 

Many calculations in linear algebra become short when the vectors involved are orthonormal. A coordinate is then a single inner product, an inverse is a transpose, and a length is a plain sum of squares. So you should know exactly what the word requires and how to get there.

Orthonormality is a characteristics of the relations among two of more vectors. It is defined by following characteristics.

  • each vectors (each column vector or row vector) should be orthogonal to each other. (It means dot product of any two vectors in the matrix should be 0).
  • Norm of  each vector should be unit length

I'll first write the two conditions as one equation and test a few small sets of vectors against it. Then we'll turn vectors that are not orthonormal into vectors that are. Finally, we'll look at how an orthonormal set simplifies a calculation.

How are the two conditions written as one equation ?

The two bullets above are easy to state in words, but a calculation needs a single test. Both conditions are statements about inner products, so one formula can hold them together.

Take vectors u1, u2, ..., un of the same size. The set is orthonormal when uiTuj = 0 for i not equal to j, and uiTui = 1 for every i. The first case is the orthogonality condition of the first bullet. The second case is the unit length condition of the second bullet, because uiTui is the squared norm of ui. The short form of both cases is uiTuj = δij, where the Kronecker delta δij is 1 for i = j and 0 otherwise.

The table below applies this test to four small sets. Only the rows where both columns say yes are orthonormal.

 

Vectors

Orthogonal ?

Unit length ?

Orthonormal ?

[1, 0] and [0, 1]

yes, 1 x 0 + 0 x 1 = 0

yes, both norms are 1

yes

[1, 1] and [1, -1]

yes, 1 - 1 = 0

no, both norms are √2

no

[1, 1]/√2 and [1, -1]/√2

yes

yes, (1/2) + (1/2) = 1

yes

[1, 0] and [1, 1]/√2

no, the inner product is 1/√2

yes

no

 

  • Both conditions are needed : the second row is orthogonal but not normalized, and the fourth row is normalized but not orthogonal.
  • Normalizing does not change orthogonality : the third row is the second row divided by √2. Scaling a vector never changes the angle between it and another vector.
  • The zero vector is never part of an orthonormal set : its norm is 0, so it cannot be scaled to length 1.
  • Complex vectors use the conjugate transpose : the test becomes uiHuj = δij. For p = [1, j]/√2 and q = [1, -j]/√2, the plain product pTq is 1, but pHq is 0. So p and q are orthonormal, and only the H form shows it.

How do you turn a set of vectors into an orthonormal set ?

Real problems rarely give you orthonormal vectors. They give you vectors that are independent but point in arbitrary directions with arbitrary lengths. So we need a procedure that keeps the space spanned by the vectors and changes only their directions and lengths.

If the vectors are already orthogonal, normalizing is the only step. Divide each vector by its norm, as in the third row of the table. If they are not orthogonal, the Gram-Schmidt procedure fixes the directions first. It takes the vectors one at a time. From each new vector, it subtracts the parts that point along the orthonormal vectors found so far, and then it normalizes what is left.

Let's run it on v1 = [1, 1, 0] and v2 = [1, 0, 1]. The first step only normalizes, so u1 = [1, 1, 0]/√2. The second step removes the part of v2 along u1. That part has the length u1Tv2 = 1/√2. So w2 = v2 - (1/√2)u1 = [1/2, -1/2, 1]. Its norm is √(3/2), and normalizing gives u2 = [1, -1, 2]/√6. A quick check confirms the result: u1Tu2 = (1 - 1 + 0)/√12 = 0, and u2Tu2 = (1 + 1 + 4)/6 = 1.

u1 = v1 / |v1|                        = [1, 1, 0] / sqrt(2)
w2 = v2 - (u1^T v2) u1                = [1/2, -1/2, 1]
u2 = w2 / |w2|                        = [1, -1, 2] / sqrt(6)
w3 = v3 - (u1^T v3) u1 - (u2^T v3) u2   (a third vector continues the same way)
  • Gram-Schmidt keeps the span : u1 and u2 span the same plane as v1 and v2. Only the description of the plane changes. See Gram-Schmidt for more examples.
  • A dependent vector gives a zero remainder : if vk lies in the span of the earlier vectors, wk is 0 and cannot be normalized. The procedure then drops it.
  • The QR decomposition is Gram-Schmidt in matrix form : the orthonormal vectors are the columns of Q, and the subtracted lengths fill R. See QR Decomposition.
  • Software uses a more stable variant : the classical procedure above loses orthogonality through rounding when the vectors are nearly parallel. Library routines such as Matlab qr and numpy.linalg.qr use Householder reflections instead.

Why is an orthonormal set so convenient ?

Now we can see the benefit of this effort. With an arbitrary basis, finding the coordinates of a vector means solving a linear system. With an orthonormal basis, each coordinate is one inner product, and no system has to be solved.

Suppose u1, ..., un is an orthonormal basis and x = c1u1 + ... + cnun. Multiply both sides by uiT. Every term except one becomes 0, and the remaining term gives ci = uiTx. For example, take u1 = [1, 1]/√2, u2 = [1, -1]/√2 and x = [3, 1]. Then c1 = 4/√2 = 2√2, which is about 2.828, and c2 = 2/√2 = √2, which is about 1.414. The length is also kept. The squared norm of x is 9 + 1 = 10, and c12 + c22 = 8 + 2 = 10.

In matrix form, put the orthonormal vectors into the columns of a matrix Q. Then QTQ = I. If Q is square, it is an orthogonal matrix, and QQT = I as well. If Q is tall, for example the 3 x 2 matrix [u1 u2] from the Gram-Schmidt example, then QQT is not I. It is the projection onto the plane spanned by u1 and u2, and in this case its diagonal is 2/3, 2/3, 2/3.

  • Coordinates are inner products : ci = uiTx, or c = QTx for all coordinates at once.
  • Lengths and energies add up : |x|2 = c12 + ... + cn2. This is Parseval's relation for an orthonormal basis.
  • QTQ = I holds for every Q with orthonormal columns : QQT = I holds only when Q is square.
  • Engineering uses it everywhere : the normalized DFT vectors, with entries e-j2πkn/N/√N, are orthonormal. The same holds for the U and V of an SVD and for the Q of a QR decomposition.