skip to content

A memory array corrects single-bit flips per word, but flips accumulate over months - how do you keep stored words recoverable?

level: principalimportance: should knowfreq 34%

answer

  1. repair happens on access, not by itself
  2. cold data collects flips unnoticed
  3. bound the window, not the rate
  4. two flips in one word is the event
  5. repeat corrections at one position mean retirement

basics

~20 s

Repair is only applied when a word is read, so rarely touched words accumulate flips until two share one word and correction fails. A background scrub that reads and rewrites every word on a fixed period bounds that accumulation window, and the corrected-error counts it produces drive retirement decisions.

solid answer

~50 s

A single-error-correcting code repairs a word **at the moment something reads it**, so a word nobody touches keeps whatever flips it has collected. Two flips in one word is the uncorrectable case, and the chance of reaching it grows with the time since the word was last repaired. A **background scrub** walks the whole array on a period, reads each word, corrects what it finds and writes the repaired value back, which caps the exposure window at the scrub period rather than the access interval. The design decisions are the period, traded against the bandwidth the scrub steals; the code strength, since correct-one-detect-two turns the double flip into a loud failure instead of silent wrong data; and the telemetry, because repeated corrections at the same position mean a permanently failed cell that rewriting cannot fix and that should be retired rather than repaired forever.

go deeper

for a junior

The fact to hold on to is that a word-level code repairs a word only when something reads it. Data nobody reads keeps its flips, and flips that pile up in one word eventually pass what the code can repair.

for a middle

Explain the mechanism of a background sweep: read every word on a period, correct, write back. Be clear about what it bounds - the accumulation window - and what it leaves untouched, which is the underlying flip rate.

for a senior

Show the operating picture: pick the period from measured correctable-error rates, distinguish transient flips from a dead cell by position history, and insist that corrections are counted so a region can be retired before it fails hard.

for a principal

Treat strength, scrub period and telemetry as three independent levers against one reliability target, only one of which costs check bits. Be explicit about the blast radius of an uncorrectable word and about who handles it above the code.

## Why correction alone is not a durability story A word-level code is a **repair on access**. Nothing in the code touches stored bits by itself: the syndrome is computed when a word is read, and the repaired value only becomes durable if it is written back. That leaves a hole that grows with time: - A hot word is read constantly, so any flip is found and repaired almost immediately. - A cold word - a rarely read region, a snapshot, anything written once - keeps every flip it has collected, invisibly. - Flips accumulate independently, so a word that already holds one correctable flip is one event away from an uncorrectable one. The failure is therefore not "a bit flipped" but "a second bit flipped in a word that was already damaged and nobody looked". Time, not access volume, is the variable that drives it. ## What scrubbing changes A **scrub** is a background sweep that reads every word on a fixed period, lets the decoder correct what it finds, and writes the repaired word back. Its effect is narrow and worth stating precisely: it does not reduce the rate at which cells flip, and it does not strengthen the code. It **bounds the window** during which flips can accumulate in one word - from "since this word was last read", which may be unbounded, to the scrub period. That matters because the uncorrectable case needs two flips in the same word within one window. Shrinking the window shrinks the exposure roughly with the square, since both flips must land inside it, which is why a scrub period is one of the highest-leverage reliability knobs available for stored data. The costs are real and belong in the decision: - **Bandwidth and power.** The sweep competes with real work; the period is chosen so the steal is a small fraction of capacity. - **Write amplification.** Writing back a repaired word touches storage that would otherwise be idle. - **No help for a word already past repair.** A scrub that finds two flips reports an uncorrectable error; it cannot undo what accumulated before it arrived. ## The three decisions a lead actually owns 1. **Scrub period.** Short enough that the probability of two flips inside one window is negligible for the flip rate observed, long enough that the sweep is invisible in the workload. This is an empirical number, derived from the measured correctable-error rate rather than assumed. 2. **Code strength and word width.** Correct-one gives silent miscorrection on a double flip; correct-one-detect-two turns the same event into a reported failure. Wider words lower the percentage overhead and raise the odds that two flips share a word. Stronger codes from other families cost more check bits and more decode work per access. 3. **What the decoder reports.** A decoder that repairs quietly throws away the only cheap predictive signal in the system. | approach | what it fixes | what it does not fix | |---|---|---| | correct on access only | flips in frequently read words | accumulation in cold words | | periodic scrub | accumulation, by bounding the window | a cell that has failed permanently | | correct-one-detect-two | silent wrong data on a double flip | the double flip itself | | retire the failing region | a cell that keeps flipping the same way | transient flips elsewhere | ## Transient versus permanent, and why the counts matter The two failure modes look identical in a single event and completely different over time: - A **transient** flip is a one-off disturbance. Rewriting the corrected value restores the word, and the position is no more likely to fail again than any other. Scrubbing handles it entirely. - A **permanent** failure is a cell that no longer holds a value. The decoder corrects it on every single read, the scrub rewrites it, and it comes back. Nothing about that is repairable by more scrubbing; the word or the region has to be retired and its contents relocated. The distinguishing evidence is the **position history**: the same position correcting repeatedly in the same word is a hard failure; scattered positions correcting occasionally across many words is background noise. That distinction only exists if corrected errors are counted and attributed, which is why "export the corrected-error rate per region" is part of the design rather than an operational afterthought. A steadily rising correctable rate in one region is a prediction of an uncorrectable event that has not happened yet, and it is the cheapest one available. ## What good judgement sounds like here The weak answer is "use a stronger code". The strong answer sets a reliability target for the data, picks the scrub period from the measured flip rate rather than from a default, chooses detect-two so that failures are loud rather than silent, insists that corrections are counted and attributed, and has a retirement path for regions whose correctable rate says a cell has died. Strength, period and telemetry are three separate levers, and only the first one costs bits.

  • Why does scrubbing not help against a cell that has failed permanently?
    Because the scrub's repair is a rewrite, and a cell that no longer holds a value returns the wrong bit on the next read regardless. The decoder corrects it every time and the error never clears. The signature is the same position correcting repeatedly in the same word, and the response is retirement and relocation, not more scrubbing.
  • How does halving the scrub period change the uncorrectable-error rate?
    More than proportionally. An uncorrectable word needs two flips inside one window, so the probability falls roughly with the square of the window length rather than linearly. Halving the period therefore cuts the rate by around a factor of four, until the residual is dominated by permanently failed cells and by multi-bit events that arrive together.
  • What is wrong with a decoder that corrects flips silently?
    It discards the only early-warning signal the system has. Corrected errors cost nothing at the time, but their rate and their position history are what distinguish background noise from a dying region and what set the scrub period empirically. Without the counts, the first evidence of trouble is an uncorrectable word, which is far too late.
  • Is a stronger code an alternative to scrubbing?
    Only partly. A stronger code raises how many flips a word survives, but the flips still accumulate without bound in data nobody reads, so a stronger code delays the event rather than preventing it. Scrubbing addresses the accumulation itself; the two are complementary levers, and scrubbing is the one that costs no check bits.

A correctable flip is a slow leak in one of many sealed vessels. Nobody notices until a second leak opens in the same vessel, so you send a patrol round on a fixed schedule - not to stop leaks starting, but to make sure two are never open in one vessel at the same time.

saying these in an interview costs you the question

  • Assumes stored words are repaired without being read
  • Says scrubbing lowers the rate at which cells flip
  • Thinks rewriting fixes a permanently failed cell
  • Treats corrected errors as noise not worth reporting
  • Answers only with a stronger code and no scrub
  • Believes wider words are unambiguously better protected