skip to content

Why can coordinates in an orthonormal basis be read off with dot products alone?

level: middleimportance: should knowfreq 42%

answer

  1. cross terms vanish
  2. unit length leaves the coefficient bare
  3. each coordinate is a projection
  4. no linear system to solve
  5. squared coordinates sum to squared length

basics

~20 s

Because the basis vectors are mutually perpendicular and of unit length, dotting a vector with one basis vector cancels every other term and leaves that coordinate directly. No system of equations has to be solved.

solid answer

~50 s

An orthonormal set satisfies `u_i . u_j = 0` for `i != j` and `u_i . u_i = 1`. Write `v = c_1 u_1 + ... + c_n u_n` and take the dot product of both sides with `u_k`: every cross term dies because those dot products are 0, and the surviving term is `c_k * 1`, so `c_k = v . u_k`. Each piece `c_k u_k` is exactly the orthogonal projection of `v` onto `u_k`. Concretely with `u_1 = (1, 1)/sqrt(2)` and `u_2 = (1, -1)/sqrt(2)`, the vector `v = (3, 1)` gives `c_1 = 4/sqrt(2) = 2*sqrt(2)` and `c_2 = 2/sqrt(2) = sqrt(2)`, and those rebuild `(2, 2) + (1, -1) = (3, 1)`. With a general non-orthogonal basis you would instead have to solve a linear system, and lengths would not be preserved; here `||v||^2 = c_1^2 + c_2^2 = 8 + 2 = 10 = 3^2 + 1^2`.

go deeper

for a junior

Be ready to state the two conditions — perpendicular pairs and unit lengths — and to apply the recipe of dotting the vector with each basis vector on a small two-dimensional example.

for a middle

Expect to reproduce the derivation on the spot: dot both sides with one basis vector, note that the cross terms are zero, and read off the coefficient. Say explicitly what the unit length is doing.

for a senior

Demonstrate why the property is worth engineering for: independent per-coordinate computation, preserved lengths and dot products, no amplification of small perturbations, and a tolerance-based check for drift.

for a principal

Own the tradeoff of building or maintaining an orthonormal representation up front versus repeatedly solving systems in a convenient but skewed one, and be able to argue when the numerical stability justifies that cost.

## The setup A set of vectors `u_1, ..., u_n` is **orthogonal** if every pair is perpendicular, meaning `u_i . u_j = 0` whenever `i != j`. It is **orthonormal** if, in addition, each vector has unit length: `u_i . u_i = ||u_i||^2 = 1`. Those two conditions can be written together as `u_i . u_j = 1` when `i = j` and `0` otherwise. Suppose such a set can express a vector `v` as a combination `v = c_1 u_1 + c_2 u_2 + ... + c_n u_n` The numbers `c_k` are the **coordinates** of `v` in that system. The question is how to find them. ## The one-line derivation Take the dot product of both sides with a particular `u_k`. The dot product distributes over the sum: `v . u_k = c_1 (u_1 . u_k) + c_2 (u_2 . u_k) + ... + c_n (u_n . u_k)` Every term with index different from `k` contains a dot product of two distinct orthonormal vectors, which is 0, so it disappears. The only survivor is `c_k (u_k . u_k) = c_k * 1`. Therefore `c_k = v . u_k` One dot product per coordinate, no elimination, no matrix inverse, and each coordinate can be computed independently of the others. Notice what this means geometrically: `c_k u_k = (v . u_k) u_k` is precisely the orthogonal projection of `v` onto the direction `u_k`. Expressing a vector in an orthonormal system is nothing more than projecting it separately onto each axis and adding the pieces back up. ## Worked example In the plane, take the rotated pair `u_1 = (1, 1) / sqrt(2)` and `u_2 = (1, -1) / sqrt(2)` Check orthonormality: `u_1 . u_2 = (1*1 + 1*(-1)) / 2 = 0`, and `u_1 . u_1 = (1 + 1)/2 = 1`, likewise for `u_2`. Good. Now take `v = (3, 1)`: - `c_1 = v . u_1 = (3 + 1)/sqrt(2) = 4/sqrt(2) = 2*sqrt(2)`, about 2.828 - `c_2 = v . u_2 = (3 - 1)/sqrt(2) = 2/sqrt(2) = sqrt(2)`, about 1.414 Rebuild to confirm: `c_1 u_1 = 2*sqrt(2) * (1,1)/sqrt(2) = (2, 2)` and `c_2 u_2 = sqrt(2) * (1,-1)/sqrt(2) = (1, -1)`. Their sum is `(3, 1)`, which is `v`. ## What orthonormality buys you **No system to solve.** For a general basis `w_1, ..., w_n` the equation `v = sum of c_k w_k` is a genuine linear system in the unknowns `c_k`; you solve it by elimination or by inverting a matrix, at cost growing like `n^3`. Orthonormality replaces that with `n` dot products. **Lengths are preserved.** Because the cross terms vanish when you expand `v . v`, you get `||v||^2 = c_1^2 + c_2^2 + ... + c_n^2` In the example, `c_1^2 + c_2^2 = 8 + 2 = 10`, and `||v||^2 = 9 + 1 = 10`. The coordinates carry the same total energy as the original components. This is the finite-dimensional form of Parseval's identity, and it fails badly for a skewed basis, where cross terms contribute. **Dot products are preserved.** For two vectors with coordinate lists `c` and `d` in the same orthonormal system, `v . w = sum of c_k d_k`. Angles and similarities computed from coordinates match those computed from the originals, so switching to an orthonormal system is a rigid motion — a rotation, possibly with a reflection — not a distortion. **Errors do not amplify.** Since the transformation preserves lengths, a small perturbation of `v` produces an equally small perturbation of the coordinates. Skewed, nearly-parallel systems can do the opposite: a tiny change in `v` swings the coordinates wildly. ## Orthogonal but not normalised If the vectors are mutually perpendicular but not of unit length, the derivation still kills the cross terms but the surviving factor is `u_k . u_k`, which is no longer 1. Then `c_k = (v . u_k) / (u_k . u_k)` which is the familiar projection coefficient. Normalising in advance is what removes that division — a good reason to store unit-length vectors when you plan to use them repeatedly. ## Practical checks To verify a candidate set is orthonormal, compute every pairwise dot product and every self dot product: off-diagonal values should be 0 and diagonal values 1, up to rounding. In floating point, expect values like `1e-16` rather than exact zeros, and judge with a tolerance relative to the sizes involved. A set whose off-diagonal entries drift to `1e-3` is not orthonormal in any useful sense, and coordinates read off with plain dot products will carry that error.

  • What changes if the vectors are mutually orthogonal but not unit length?
    The cross terms still vanish, but the surviving factor is `u_k . u_k` instead of 1, so the coordinate becomes `c_k = (v . u_k) / (u_k . u_k)` — the ordinary projection coefficient. You pay one extra division per coordinate. Normalising the set up front removes it, which is why orthonormal is the preferred form when the vectors are reused.
  • Does expressing a vector in an orthonormal system change its length?
    No. Expanding `v . v` with orthonormal vectors kills every cross term and leaves `||v||^2 = c_1^2 + ... + c_n^2`. Dot products between two vectors are preserved too, so lengths and angles all survive intact — the change of coordinates is a rotation or reflection, not a stretch or a skew.
  • How would you check that a given set of vectors is orthonormal?
    Compute all pairwise dot products and all self dot products: distinct pairs should give 0 and each vector with itself should give 1. In floating point, compare against a tolerance rather than testing for exact equality, since results like `1e-16` are normal. Drift of `1e-3` off the diagonal means the set has lost orthogonality in practice.

saying these in an interview costs you the question

  • Solves a linear system when the vectors are already orthonormal
  • Confuses orthogonal with orthonormal and skips the unit lengths
  • Thinks coordinates must be computed jointly rather than one at a time
  • Claims a change of orthonormal coordinates rescales the vector
  • Tests floating-point orthogonality with exact equality to zero

context