Two lossy previews show the same mean squared error against the original; why can one still look far worse?
answer
- it counts samples, not structure
- where the error sits matters
- texture masks, flat areas expose
- decibels restate the same number
- chosen for tractability, not for vision
basics
~20 sMean squared error averages per-sample differences and ignores where they land. The same total spread thinly across busy texture is invisible; concentrated in a flat region it is obvious. Peak signal-to-noise ratio only restates that number in decibels.
solid answer
~40 s**Mean squared error** is the average of the squared differences sample by sample. It has no model of a viewer: it does not care whether the error sits in a smooth sky or a patch of foliage, whether it forms a visible edge, or whether it is a constant offset nobody notices. Perception does care — busy texture **masks** error, flat areas expose it, and structured artifacts read as damage while the same energy spread as noise does not. **Peak signal-to-noise ratio** is `10 * log10(MAX^2 / MSE)`, a monotone re-expression of the same number in decibels, so it inherits every blind spot. Squared error survives because it is additive, cheap and gives closed-form rate-distortion results — not because it models vision.
go deeper
Know that squared error is an average of per-sample differences and that the decibel figure is the same number rescaled, so neither one knows anything about how the picture looks.
Explain masking and concentration: the same total error is invisible spread across texture and obvious in a flat gradient, and a shift or offset scores terribly while looking fine.
Use the measures for what they are good at — regression detection — and insist on a structural measure plus a sampled viewer check before an operating point becomes a policy.
Decide what the organisation actually optimises. A cheap additive measure buys tractable allocation; a perceptual one buys relevance at the cost of optimisability, and the choice moves the whole trade-off curve.
## What squared error counts For an original `x` and a reconstruction `y` of the same length `n`, **mean squared error** is `MSE = (1/n) * sum over i of (x[i] - y[i])^2`. Three properties follow directly, and each is a blind spot: - It is **per-sample**. Each difference is scored on its own, so the measure cannot see that a set of differences forms an edge, a block boundary or a ring around a contour. - It is **position-blind**. Moving the same error from a textured region to a flat one changes nothing in the number and changes a great deal on screen. - It is **symmetric and unweighted**. A difference that a viewer cannot resolve counts exactly as much as one that jumps out. **Peak signal-to-noise ratio** adds no information: `PSNR = 10 * log10(MAX^2 / MSE)` where `MAX` is the largest representable sample value. It is a decibel re-scaling, useful because it compresses a wide range into readable numbers and because differences in decibels are easy to compare — but it is a strictly decreasing function of `MSE`, so equal squared error means equal PSNR, and every blind spot carries over. ## Where the measure and the eye disagree | Situation | Squared error | What a viewer sees | |---|---|---| | Small error spread over dense texture | Meaningful | Nothing — texture masks it | | The same total concentrated in a smooth gradient | Identical | Obvious banding or blotches | | A small constant brightness offset everywhere | Large | Almost nothing | | A one-pixel shift of the whole frame | Very large | Indistinguishable | | A thin structured artifact along a strong edge | Small | Immediately visible as damage | The last two rows are the ones that surprise people. A shift or an offset leaves every structure intact and every sample wrong, which is the worst case for a per-sample measure and a non-event for a viewer. Conversely a **structured** error of tiny energy — a halo on a face, a discontinuity across a block boundary in an otherwise smooth area — reads instantly as damage, because vision is built to find structure. There is a second, quieter trap: decibel figures are **not comparable across content**. The same value on a noisy photograph and on a flat graphic describe completely different viewing experiences, because the amount of detail available to mask error differs. A dashboard that averages decibels over a heterogeneous library produces a number that means very little. ## Why the theory still uses squared error If it models perception so poorly, why is it everywhere? 1. **It is additive.** Total error decomposes over samples, regions and transform components, which is what makes allocation arguments tractable. 2. **It is differentiable and cheap**, so it can sit inside an encoder's inner loop and be optimised directly. 3. **It has closed-form results.** The clean rate-distortion statements — the ones that give a rate as an explicit function of a distortion budget — are stated under squared error. 4. **The theory is agnostic to the measure.** The rate-distortion framework takes *some* distortion function as given. Squared error is a choice inside it, not part of it. Swapping in a perceptual measure is legitimate and changes the curve and the chosen operating point; what it costs is the additivity and the closed forms. ## What to do about it in practice - **Use squared error or decibels as a regression guard, not as a quality target.** They are excellent at telling you that something changed and poor at telling you whether anyone minds. - **Prefer a measure that scores structure** when you are choosing an operating point, and accept that it will be more expensive and harder to optimise against. Structural-similarity-style measures compare local statistics rather than individual samples, which is why they catch banding and halos that squared error waves through. - **Validate on a sample with real viewers.** The distortion measure is a proxy for a judgement, and which proxy tracks the judgement is a property of your content, not a universal fact. - **Never compare decibel figures across different content** and never average them over a mixed library without saying what the average is over. The short version for an interview: squared error measures how far the numbers moved; perception measures how much the structure broke. They agree often enough to be useful and disagree exactly where lossy encoding puts its artifacts.
- Does choosing a different distortion measure change the rate-distortion curve?Yes. The curve is defined relative to a chosen distortion function and budget, so swapping the measure moves the curve and the operating point you would pick from it. The structural properties survive — it is still non-increasing and convex in the budget — but the numbers are not comparable between measures, and neither are the quality settings chosen under them.
- Why does a tiny structured artifact score better than a large invisible offset?Because squared error sums per-sample differences with no notion of structure. A constant offset makes every sample wrong by a little, which is a large sum and an invisible change. A halo along one edge makes few samples wrong, which is a small sum, but vision is tuned to find exactly that kind of local structure.
It is like grading an essay by counting changed characters. A rewritten sentence and a global find-and-replace of one letter can score the same, and only one of them changes what the reader takes away.
saying these in an interview costs you the question
- Reads peak signal-to-noise ratio as a perceptual score.
- Compares decibel figures across completely different content.
- Assumes equal squared error means equal viewer experience.
- Thinks squared error is used because it models human vision.
- Believes a small global brightness shift barely moves squared error.
- Treats a decibel target as a product quality requirement.