skip to content

Why is a tiny determinant not by itself evidence that a matrix is nearly singular?

level: seniorimportance: nice to knowfreq 28%

answer

  1. the determinant is not unit-free
  2. rescaling every entry moves it
  3. the exponent is the dimension
  4. one direction can cancel another
  5. compare most-stretched to least-stretched

basics

~10 s

Because the determinant depends on scale: for an n x n matrix, det(cA) = c^n det(A). Shrinking every entry drives the determinant toward zero without making the matrix any harder to invert.

solid answer

~50 s

The determinant is a volume scaling factor, so it inherits the units and the scale of the matrix and it compounds with dimension. Take `0.01 * I` at size 10: its determinant is `10^-20`, yet the matrix is a plain scalar multiple of the identity whose inverse is `100 * I` — nothing about it is delicate. The converse fails too: `[[10^6, 0], [0, 10^-6]]` has determinant exactly 1, while it stretches one direction by a million and crushes the other by the same factor, which is far closer to flattening than the first example. So a determinant near zero can mean "small numbers" and a determinant near 1 can hide extreme distortion. `det(A) = 0` is an exact, meaningful statement in exact arithmetic; `det(A) is small` is not a calibrated measure of anything until you say small relative to what.

go deeper

for a junior

Take away one fact: a determinant close to zero is not the same as a determinant of zero, and only the exact zero decides invertibility.

for a middle

Explain the scaling rule det(cA) equals c to the n times det(A) and use it to show that shrinking every entry drives the determinant toward zero while changing nothing about the map.

for a senior

Show the failure in both directions with concrete matrices, and offer the scale-free alternative of comparing the most-stretched with the least-stretched direction.

for a principal

Own the standard: refuse thresholds on scale-dependent quantities, insist that any numerical health check be invariant to units, and say who decides the tolerance and on what evidence.

## The trap The theorem people remember is clean: a square matrix is invertible exactly when its determinant is non-zero. The tempting extrapolation is that the *magnitude* of the determinant grades how invertible a matrix is — big determinant, comfortable; small determinant, nearly singular. That extrapolation is wrong, and interviewers use it to separate people who memorised the theorem from people who understand what the determinant measures. ## Reason one: the determinant carries scale Multiplying every entry of an n-by-n matrix by a constant `c` multiplies the determinant by `c^n`, because volume scaling compounds across all n directions: ``` det(cA) = c^n * det(A) ``` Consider `A = 0.01 * I` with `I` the 10-by-10 identity. Then `det(A) = 0.01^10 = 10^-20`. By the "small determinant means nearly singular" heuristic this matrix would be on the brink of collapse. In fact it is the tamest matrix imaginable: its inverse is `100 * I`, exactly, and it treats all directions identically. All that changed was the unit in which the entries are expressed — swap metres for centimetres in a physical problem and every determinant in sight moves by orders of magnitude while nothing about the underlying map changes. The dimension exponent makes this worse as matrices grow. At n = 100, a uniform factor of `0.5` moves the determinant by `2^-100`. Any threshold like "call it singular below `10^-12`" is therefore a statement about the size and units of your entries, not about the geometry of the map. ## Reason two: the determinant aggregates and cancels The determinant is a single product-like aggregate over all directions, so expansion in one direction can cancel compression in another. Take ``` B = [[10^6, 0], [0, 10^-6]] ``` Its determinant is `10^6 * 10^-6 = 1` — the number you would associate with a perfectly well-behaved, volume-preserving map. Yet `B` stretches the horizontal direction by a million and crushes the vertical by a million. It maps the unit square to an extremely long, extremely thin sliver: visually almost flat, and far closer to a collapse than `0.01 * I` ever was. Undoing it amplifies anything that happened in the crushed direction by `10^6`. So the two examples run in opposite directions to the naive heuristic. The one with determinant `10^-20` is benign; the one with determinant 1 is delicate. ## What actually characterises nearness to singularity A collapse is a direction being squashed to nothing, so a scale-free way to ask "how close to flat is this map?" compares the most-stretched direction against the least-stretched one. If the least-stretched direction is tiny relative to the most-stretched, the map is nearly flattening along that direction, whatever the determinant happens to be. That ratio is unchanged when you rescale the whole matrix — rescaling multiplies both ends by the same factor — which is exactly the invariance the determinant lacks. If you insist on using the determinant, normalise it first, for example by dividing by the product of the row norms, so that changing units cannot move the answer. ## The floating-point side Even when the determinant is the right quantity, computing it for a large matrix is awkward, because it is effectively a product of n numbers. With n in the hundreds, a product of factors that are individually unremarkable will overflow or underflow double-precision range, returning `0` or infinity for a matrix that is neither. The standard dodge is to work with `log|det(A)|`, accumulating a sum of logarithms instead of a product, and to keep the sign separately. That keeps the quantity representable, but it does not fix the scale-dependence problem — a log of a scale-dependent quantity is still scale-dependent. ## When the determinant is genuinely the right tool None of this makes the determinant useless. It is exactly right as (1) an *exact* zero-or-not test in exact arithmetic or on small symbolic examples, (2) a volume scaling factor, which is a real quantity people need when a change of variables rescales volumes, and (3) a hand-computable check for 2-by-2 and 3-by-3 matrices. The mistake is only the extra step: turning a magnitude that has no fixed scale into a graded diagnosis. ## Answering in the room Give `det(cA) = c^n det(A)` and the `0.01 * I` example, then flip it with the determinant-1 sliver so the interviewer sees you know the failure runs both ways. Close with the correction: nearness to singularity is about the ratio between the most- and least-stretched directions, which is scale-free, and if you must use a determinant, normalise it and be explicit about the units it lives in.

  • If A is 5x5 with det(A) = 3, what is det(0.1 * A)?
    It is `0.1^5 * 3 = 0.00003`. Scaling an n-by-n matrix by `c` scales the determinant by `c^n`, so a modest rescaling of the entries moves the determinant by five orders of magnitude here. That exponent is precisely why a fixed numeric threshold on the determinant is meaningless without knowing the size and units of the entries.
  • Why do people work with the logarithm of the absolute determinant for large matrices?
    Because a determinant behaves like a product of n numbers, so for n in the hundreds it overflows or underflows floating-point range and returns 0 or infinity for a matrix that is neither. Accumulating `log|det|` turns the product into a sum of logarithms, which stays representable; the sign is tracked separately.
  • So when is the determinant genuinely the right quantity to compute?
    When you need an exact zero-or-not answer in exact arithmetic, when you need the volume scaling factor itself — for example because a change of variables rescales volumes — or when the matrix is small enough to compute by hand. It is a real quantity with a real meaning; the mistake is reading its magnitude as a graded measure of invertibility.

It is like judging whether a box is nearly flat from its volume alone: a matchbox has a small volume but is not flat, while a very wide, very thin sheet can have the same volume as a cube and be almost flat.

saying these in an interview costs you the question

  • Treats determinant magnitude as how invertible a matrix is
  • Forgets that det(cA) equals c to the n times det(A)
  • Proposes a fixed numeric threshold for calling a matrix singular
  • Assumes determinant 1 means the matrix is well behaved
  • Believes a computed determinant of zero always means singular

context