A radio link measures a 0.1 percent bit error rate, yet whole blocks fail far more often than an independent-flip model predicts; why?
answer
- the average survived, independence did not
- same mean, different distribution
- more clean blocks and more dead blocks
- measure errors per block, not the mean
- two-state good and bad regimes
basics
~20 sThe average rate is right but the memoryless assumption is wrong. Fading concentrates errors into bursts, so most blocks arrive perfectly clean while the few blocks a burst touches carry far more errors than any fixed repair budget covers.
solid answer
~40 sA bit error rate is an **average**; a memoryless model additionally assumes errors are **independent**, and that second claim is what fading breaks. Take 0.1 percent over thousand-bit blocks. Independently, each block averages one flip, 63 percent of blocks are touched, and with a budget of five repairable errors per block the residual failure rate is about 0.06 percent. Now concentrate the same average into bursts: one block in a hundred takes about 100 flips and the rest are untouched. Now 99 percent of blocks are **cleaner** than before, yet every affected block blows through a five-error budget, so the block failure rate is 1 percent — roughly seventeen times worse. Same average, opposite conclusion. Measure the distribution of errors per block, not just the mean.
code
pseudocode · 13 linesstate = GOOD
for i in 0 .. n-1:
if state == GOOD:
flipped[i] = (random() < p_good) # p_good is tiny
if random() < enter_bad:
state = BAD
else:
flipped[i] = (random() < p_bad) # p_bad is large
if random() < leave_bad:
state = GOOD
# average rate depends on p_good, p_bad and the
# fraction of time spent in the BAD statego deeper
The takeaway is that an average error rate does not say how the errors are spread out. Errors that arrive in clumps behave very differently from errors sprinkled evenly, even when the average matches.
Explain the memoryless assumption and why it is separate from the average. Be able to compute the clean-block probability under independence and say why a clustered link departs from it in both directions.
Show the diagnostic instinct: ask for errors per block and run lengths, compare the observed histogram against the independent prediction, and connect the tail past the repair budget to the failure rate operators actually see.
The trade-off to own is how much modelling fidelity a reliability target justifies. A four-parameter two-regime model costs measurement and complexity; committing to a one-parameter model instead is a cheaper bet with a known failure mode that should be stated up front.
## The assumption that fails, and the one that does not A link characterised as "0.1 percent bit error rate" has been summarised by a single average. A **memoryless** channel model adds a second claim on top of that average: that each bit's fate is **independent** of its neighbours'. Physical impairments routinely violate the second claim while respecting the first. A fade, an interferer, a mechanical defect on a stored medium — each one produces a run of consecutive damaged bits and then nothing for a long while. The average survives; the distribution does not. This matters because almost every downstream decision is made against the *distribution*, not the average. A repair scheme has a budget — some number of damaged symbols per block it can absorb — and what you need to know is how often a block exceeds that budget. ## Same average, opposite conclusion Take thousand-bit blocks and an average bit error rate of 0.001, and give the system a budget of five repairable errors per block. **Independent model.** Errors per block are Poisson with mean 1. Then: - 36.8 percent of blocks arrive with no error at all; - 63.2 percent contain at least one error, nearly all of them one or two; - the chance of six or more errors — the budget being exceeded — is about 0.00059, so roughly **6 blocks in 10,000 fail**. **Bursty model with the same average.** Suppose a fade hits one block in a hundred and corrupts about 100 bits within it. The average checks out: `0.01 * 100 / 1000 = 0.001`. Then: - 99 percent of blocks arrive with no error at all; - 1 percent of blocks contain about 100 errors; - every one of those blows through the five-error budget, so **100 blocks in 10,000 fail**. | statistic | independent flips | bursty, same average | |---|---|---| | error-free blocks | 36.8% | 99% | | blocks containing any error | 63.2% | 1% | | typical errors in an affected block | 1 to 2 | about 100 | | blocks exceeding a five-error budget | 0.06% | 1% | ## Which statistic moves which way This is where careless reasoning shows. Bursts do **not** simply make everything worse: - The **error-free block fraction rises**, often dramatically — from 37 percent to 99 percent here. A dashboard tracking "percentage of clean blocks" would show a bursty link looking *better*. - The **failure rate against a fixed repair budget rises** — by about seventeen times in this example — because damage arrives in lumps too big for the budget. - The **average bit error rate is unchanged**, which is precisely why it is a poor predictor on its own. Whether bursts help or hurt therefore depends on what you measure and what budget you hold. Against a very generous budget, clustering can even be favourable: fewer blocks are touched at all. Against a small budget, clustering is punishing. Stating the direction without naming the budget is the mistake to avoid. ## What to measure instead 1. **Errors per block, as a histogram.** The mean tells you almost nothing; the tail past your budget is the number that predicts failures. 2. **Run lengths.** How long does a bad period last, in bits or in time? That length compared to your block size decides whether damage lands inside one block or spreads over several. 3. **The gap distribution between bad periods.** This is what tells you whether the link has two regimes rather than one. A useful modelling shape is a **two-state** channel: a good state with a tiny flip probability and a bad state with a large one, plus transition probabilities between them. It reproduces both the average and the clustering with four parameters instead of one, and it is simple enough to simulate on a whiteboard. The standard engineering response to bursts — reordering symbols so that consecutive damaged bits land in different blocks — is a coding-side subject and belongs with symbol block codes, not with the channel model itself. ## What interviewers listen for - That you separate the average from the distribution and say which one the design actually depends on. - That you name the memoryless assumption explicitly as the thing being violated, rather than blaming "noise". - That you get the directions right: more clean blocks, more unrecoverable blocks, unchanged average. - That you propose measuring errors per block and run lengths, instead of re-deriving from a single rate.
- Does making a link's errors bursty always raise its block failure rate?No, and the direction depends on the repair budget. Against a small budget, clustering is punishing, because a single burst exceeds it while scattered errors would not. Against a budget larger than a typical burst, clustering can be favourable: the same total damage is confined to fewer blocks, and every other block arrives untouched. State the budget before stating the direction.
- How would you tell from measurements alone that a link is bursty rather than memoryless?Collect errors per block as a histogram and compare it with the Poisson distribution implied by the measured average. A bursty link shows a spike at zero far above the independent prediction and a long tail of heavily damaged blocks, with the middle of the distribution almost empty. Run-length statistics confirm it: consecutive damaged bits far longer than chance would give.
- Why is the average bit error rate still worth measuring if it predicts block failures so poorly?It is a cheap, stable summary of the physical link's health and a good trending signal: a rising average almost always means something physical changed. What it cannot do is predict failures against a repair budget on its own, because two links with identical averages can differ by more than an order of magnitude in block failure rate.
saying these in an interview costs you the question
- Claims bursts always reduce the fraction of clean blocks
- Treats the average bit error rate as the whole model
- Says an average error rate implies independent errors
- States bursts are worse without naming any repair budget
- Assumes measured averages already capture clustering
- Believes a memoryless model can be fixed by raising p