skip to content

Two lossy previews show the same mean squared error against the original; why can one still look far worse?

level: middleimportance: nice to knowfreq 30%

answer

  1. it counts samples, not structure
  2. where the error sits matters
  3. texture masks, flat areas expose
  4. decibels restate the same number
  5. chosen for tractability, not for vision

basics

~20 s

Mean squared error averages per-sample differences and ignores where they land. The same total spread thinly across busy texture is invisible; concentrated in a flat region it is obvious. Peak signal-to-noise ratio only restates that number in decibels.

solid answer

~40 s

**Mean squared error** is the average of the squared differences sample by sample. It has no model of a viewer: it does not care whether the error sits in a smooth sky or a patch of foliage, whether it forms a visible edge, or whether it is a constant offset nobody notices. Perception does care — busy texture **masks** error, flat areas expose it, and structured artifacts read as damage while the same energy spread as noise does not. **Peak signal-to-noise ratio** is `10 * log10(MAX^2 / MSE)`, a monotone re-expression of the same number in decibels, so it inherits every blind spot. Squared error survives because it is additive, cheap and gives closed-form rate-distortion results — not because it models vision.

go deeper

for a junior

Know that squared error is an average of per-sample differences and that the decibel figure is the same number rescaled, so neither one knows anything about how the picture looks.

for a middle

Explain masking and concentration: the same total error is invisible spread across texture and obvious in a flat gradient, and a shift or offset scores terribly while looking fine.

for a senior

Use the measures for what they are good at — regression detection — and insist on a structural measure plus a sampled viewer check before an operating point becomes a policy.

for a principal

Decide what the organisation actually optimises. A cheap additive measure buys tractable allocation; a perceptual one buys relevance at the cost of optimisability, and the choice moves the whole trade-off curve.

## What squared error counts For an original `x` and a reconstruction `y` of the same length `n`, **mean squared error** is `MSE = (1/n) * sum over i of (x[i] - y[i])^2`. Three properties follow directly, and each is a blind spot: - It is **per-sample**. Each difference is scored on its own, so the measure cannot see that a set of differences forms an edge, a block boundary or a ring around a contour. - It is **position-blind**. Moving the same error from a textured region to a flat one changes nothing in the number and changes a great deal on screen. - It is **symmetric and unweighted**. A difference that a viewer cannot resolve counts exactly as much as one that jumps out. **Peak signal-to-noise ratio** adds no information: `PSNR = 10 * log10(MAX^2 / MSE)` where `MAX` is the largest representable sample value. It is a decibel re-scaling, useful because it compresses a wide range into readable numbers and because differences in decibels are easy to compare — but it is a strictly decreasing function of `MSE`, so equal squared error means equal PSNR, and every blind spot carries over. ## Where the measure and the eye disagree | Situation | Squared error | What a viewer sees | |---|---|---| | Small error spread over dense texture | Meaningful | Nothing — texture masks it | | The same total concentrated in a smooth gradient | Identical | Obvious banding or blotches | | A small constant brightness offset everywhere | Large | Almost nothing | | A one-pixel shift of the whole frame | Very large | Indistinguishable | | A thin structured artifact along a strong edge | Small | Immediately visible as damage | The last two rows are the ones that surprise people. A shift or an offset leaves every structure intact and every sample wrong, which is the worst case for a per-sample measure and a non-event for a viewer. Conversely a **structured** error of tiny energy — a halo on a face, a discontinuity across a block boundary in an otherwise smooth area — reads instantly as damage, because vision is built to find structure. There is a second, quieter trap: decibel figures are **not comparable across content**. The same value on a noisy photograph and on a flat graphic describe completely different viewing experiences, because the amount of detail available to mask error differs. A dashboard that averages decibels over a heterogeneous library produces a number that means very little. ## Why the theory still uses squared error If it models perception so poorly, why is it everywhere? 1. **It is additive.** Total error decomposes over samples, regions and transform components, which is what makes allocation arguments tractable. 2. **It is differentiable and cheap**, so it can sit inside an encoder's inner loop and be optimised directly. 3. **It has closed-form results.** The clean rate-distortion statements — the ones that give a rate as an explicit function of a distortion budget — are stated under squared error. 4. **The theory is agnostic to the measure.** The rate-distortion framework takes *some* distortion function as given. Squared error is a choice inside it, not part of it. Swapping in a perceptual measure is legitimate and changes the curve and the chosen operating point; what it costs is the additivity and the closed forms. ## What to do about it in practice - **Use squared error or decibels as a regression guard, not as a quality target.** They are excellent at telling you that something changed and poor at telling you whether anyone minds. - **Prefer a measure that scores structure** when you are choosing an operating point, and accept that it will be more expensive and harder to optimise against. Structural-similarity-style measures compare local statistics rather than individual samples, which is why they catch banding and halos that squared error waves through. - **Validate on a sample with real viewers.** The distortion measure is a proxy for a judgement, and which proxy tracks the judgement is a property of your content, not a universal fact. - **Never compare decibel figures across different content** and never average them over a mixed library without saying what the average is over. The short version for an interview: squared error measures how far the numbers moved; perception measures how much the structure broke. They agree often enough to be useful and disagree exactly where lossy encoding puts its artifacts.

  • Does choosing a different distortion measure change the rate-distortion curve?
    Yes. The curve is defined relative to a chosen distortion function and budget, so swapping the measure moves the curve and the operating point you would pick from it. The structural properties survive — it is still non-increasing and convex in the budget — but the numbers are not comparable between measures, and neither are the quality settings chosen under them.
  • Why does a tiny structured artifact score better than a large invisible offset?
    Because squared error sums per-sample differences with no notion of structure. A constant offset makes every sample wrong by a little, which is a large sum and an invisible change. A halo along one edge makes few samples wrong, which is a small sum, but vision is tuned to find exactly that kind of local structure.

It is like grading an essay by counting changed characters. A rewritten sentence and a global find-and-replace of one letter can score the same, and only one of them changes what the reader takes away.

saying these in an interview costs you the question

  • Reads peak signal-to-noise ratio as a perceptual score.
  • Compares decibel figures across completely different content.
  • Assumes equal squared error means equal viewer experience.
  • Thinks squared error is used because it models human vision.
  • Believes a small global brightness shift barely moves squared error.
  • Treats a decibel target as a product quality requirement.