Why can adding a small random signal before rounding samples to a coarse grid improve the result, despite raising total error power?
answer
- the noise model has a precondition
- small or smooth signals break independence
- error that tracks the signal is visible
- about one step of added randomness
- louder floor, no pattern
basics
~20 sUndithered rounding leaves an error that tracks the signal, so quiet or smoothly varying passages get patterned artifacts such as tonal distortion or visible banding. Adding about one step of noise decorrelates the error, trading structure for a slightly higher but featureless floor.
solid answer
~50 sThe usual model - quantization error as independent noise spread evenly across the bin - only holds when the signal crosses many levels between samples. For a signal that barely exceeds one step, the error is a deterministic function of the signal, so it arrives as harmonics and patterns rather than as noise: tonal artifacts in a fading tail, contour bands across a smooth gradient. Adding a small random offset before rounding breaks that dependence. The usual choice is triangular noise about two steps peak to peak, formed by summing two independent uniform draws, which makes both the mean and the variance of the error independent of the signal. The price is real: total error power rises to roughly three times the undithered figure, about 5 dB. You pay it because a slightly louder floor with no structure is far less objectionable than a quieter one that correlates with the content.
code
pseudocode · 12 linesstep = range / 2^b
# undithered: the error is a fixed function of x
out = round(x / step) * step
# dithered: two independent uniform draws, summed
# the sum is triangular and spans two steps peak to peak
d = uniform(-step/2, step/2) + uniform(-step/2, step/2)
out = round((x + d) / step) * step
# d is NOT subtracted afterwards: it stays in the output
# as part of the now signal-independent error floorgo deeper
Know that rounding a very quiet or smoothly varying signal produces patterns rather than hiss, and that a deliberate pinch of noise before the rounding is the standard cure.
Explain the precondition behind the familiar noise model: the error only behaves like independent noise when the signal crosses several levels between samples.
Place it correctly - once, immediately before the final rounding - and defend paying a few decibels of extra floor to get an unstructured one instead of a quieter correlated one.
Decide where structured error is tolerable at all: for perceived media it rarely is, while for measurement or financial data a documented rounding rule beats injected randomness.
## The noise model has a precondition Every convenient statement about quantization error - that it is uniform across the bin, zero-mean, independent of the signal, with power equal to the square of the step over twelve - is a **model**, and the model has a condition attached: the signal must move across several levels between consecutive samples so that successive errors are effectively unrelated. Where that holds, the error behaves like a low, featureless floor. Where it does not hold, the error is what it always literally was: a **deterministic function of the input**. Rounding is not random. Feed it a signal that spans one or two steps, or one that changes very slowly, and the error becomes a repeating, signal-shaped pattern. ## What correlated error looks like The symptoms differ by medium but the mechanism is identical: - a **fading tail** quantized to few levels breaks into a stepped staircase, heard as tonal artifacts rather than hiss, and the artifacts change pitch with the signal; - a **smooth gradient** rounded to a coarse value grid shows **contour bands**, because a whole region of nearly equal values collapses onto one level and the boundary between levels becomes a visible edge; - a **slowly drifting measurement** rounded to a coarse grid produces long runs of identical values and then a jump, which downstream statistics read as a step change that never happened. In all three, the total error power may be small. The problem is that it is **organised**, and organised error is far more noticeable than unorganised error of the same size. ## What dither actually does Dither adds a small random value to each sample **before** the rounding, and does not subtract it afterwards. The randomness moves the rounding decision off the signal's own value, so which level a given input lands on stops being a fixed function of that input. | | Undithered rounding | With triangular dither | |---|---|---| | Error vs signal | deterministic, correlated | independent in mean and variance | | Character | harmonics, bands, staircases | featureless floor | | Total error power | the square of the step over twelve | about three times that, roughly 5 dB more | | Low-level detail | collapses onto one level | survives as a modulation of the floor | The last row is the part that surprises people. A signal smaller than one step is not simply lost under dither: it shifts the *proportion* of samples rounding up versus down, so information below the step survives in the statistics of the output even though no single sample resolves it. The choice of distribution matters: 1. **One uniform draw spanning one step** makes the error's average independent of the signal, but its variance still moves with the signal, which is perceived as the floor pumping in time with the content. 2. **Two independent uniform draws, summed** - a triangular distribution spanning two steps - makes the second moment independent too, which is why it is the usual default. 3. **Anything much larger** simply adds noise without buying further decorrelation, so more is not safer. ## Where it belongs in a chain Dither goes **immediately before the final, irreversible rounding onto the delivery grid, and only there**. Applied earlier, it is re-rounded by later stages, which both re-correlates the error and stacks another noise contribution on top of the first. Each irreversible rounding earns at most one dither of its own, which makes dither placement another argument for a chain that rounds once, late. It is also not universal. It is a technique for data that is **perceived** - waveforms, images, anything a human judges - where structured error is worse than louder error. For a measurement pipeline, a counter, or a currency amount, structure in the error is not the enemy and a documented rounding rule is worth more than injected randomness; auditability beats decorrelation there. ## The answer an interviewer wants State the trade in one line: dither raises total error power in exchange for removing correlation between the error and the signal, and correlation is what is actually noticeable. Then name the precondition that makes it necessary - a signal small or smooth relative to the step - and the placement rule that keeps it from being applied several times over.
- When is dither pointless or harmful?When the signal already crosses many levels between samples, the error is effectively decorrelated already and dither only adds power. It is also the wrong tool for data that is not perceived as a waveform or an image: rounding a counter, a currency amount or a billing quantity wants a stated, auditable rounding rule, not injected randomness.
- Why triangular rather than a single uniform draw?A single uniform draw spanning one step makes the error's average independent of the signal, but its variance still moves with the signal, which is perceived as the noise floor pumping in time with the content. Summing two independent draws makes the second moment independent as well, at the cost of more added power.
- Where in a chain should dither be applied?Immediately before the final rounding onto the delivery grid, and only there. Dither added earlier is re-rounded by later stages, which re-correlates the error and stacks a second noise contribution on the first. Each irreversible rounding earns at most one dither, which is another reason to round once and late.
saying these in an interview costs you the question
- Says adding noise always improves a quantized signal
- Claims dither lowers the total error power
- Treats quantization error as noise-like for signals near one step
- Applies dither before stages that will round again
- Believes dither adds resolution rather than restructuring the error
- Uses dither far larger than a step to be safe