Rebuilding one lost fragment of a ten-of-fourteen erasure-coded blob costs far more than replacing a lost replica — why?
answer
- nothing on disk equals the missing piece
- repair recomputes rather than copies
- ten reads to restore one tenth
- amplification equals k
- domain loss multiplies it fleet-wide
basics
~20 sNothing on disk equals the missing fragment, so it must be recomputed: the repair reads ten surviving fragments, a whole blob's worth of traffic, to restore a tenth of a blob. Replication restores a lost copy by copying one copy.
solid answer
~50 sUnder replication every byte exists somewhere else verbatim, so repair is a copy: one blob read to restore one blob. Under a ten-of-fourteen code no single survivor contains the lost fragment's bytes, so the system must fetch `k` = 10 fragments and reconstruct. Each fragment is a tenth of the blob, so ten of them is one whole blob of reads and network transfer to regenerate one tenth — a tenfold amplification per repaired byte. That cost scales with the loss: a whole domain going away triggers a rebuild for every blob that had a fragment there, and the resulting traffic competes with live reads. Meanwhile any read of the affected range is a degraded read, reconstructing on the fly. Operators respond with a repair delay for transient absences, bandwidth throttles, and prioritising the blobs closest to losing durability.
code
pseudocode · 12 lineson fragment_lost(blob, lost_index):
healthy = fragments_present(blob)
if count(healthy) < k:
give_up(blob) // past the parity budget
sources = pick_any(k, healthy) // k reads, not one
for each s in sources:
bytes[s] = fetch(s) // each is blob_size / k
rebuilt = interpolate(bytes, target = lost_index)
place(rebuilt, in = domain_holding_no_fragment_of(blob))
// read traffic = k * (blob_size / k) = blob_size
// bytes restored = blob_size / kgo deeper
The key idea is that a coded fragment cannot be copied from anywhere, because no survivor holds its bytes; it has to be recomputed from several other fragments.
Quantify it: rebuilding one fragment reads k fragments, which totals one whole blob of traffic to restore a tenth of a blob, against a one-to-one copy under replication.
Talk about the fleet. A lost domain multiplies that cost across every blob it held, repair competes with live reads, and degraded reads make the rebuild window visible in latency until it closes.
The lever to own is the balance between overhead and repair load: a wider code saves storage and raises repair amplification, while local parity groups buy cheap common-case repair back with storage.
## The arithmetic of a single repair Take a blob of size `S` stored as ten data plus four parity fragments. Each fragment is `S/10`. One fragment is lost. Nothing among the survivors holds those bytes. The code's whole premise is that fragments are *different* functions of the data, not copies of it, so the missing one has to be recomputed: - fetch any `k` = 10 surviving fragments: `10 * S/10` = `S` bytes read and transferred; - interpolate the missing position and write it: `S/10` bytes restored. That is a **tenfold amplification**: a whole blob of network and disk work per tenth of a blob recovered. Under three-way replication the same loss is repaired by streaming one surviving copy: `S` read to restore `S`, an amplification of one. The ratio is `k`, and it is the direct price of the storage saving — widening the code to lower overhead raises this number in lockstep. ## Why a domain failure is a traffic event, not a data event A single fragment loss is nothing. The operational problem is that a failure domain holds fragments belonging to an enormous number of distinct blobs, and **each blob's repair is independent and pays the full `k` multiplier**. Losing one domain's worth of capacity `C` therefore generates on the order of `k * C` bytes of reads and cross-network transfer, spread over every surviving domain that holds a peer fragment. That traffic does not run in a vacuum. It shares network, disks and request budget with live reads, which is how a durability event becomes a latency incident. The failure mode to name is the cascade: repair traffic saturates the fleet, saturation slows or times out ordinary reads, timeouts are mistaken for further absences, more repairs are queued. ## Degraded reads during the rebuild window While a fragment is absent, a read that needs its byte range cannot simply be served. The reader fetches enough surviving fragments and reconstructs on the fly — a **degraded read**. It returns correct data, and it costs extra fan-out, extra bytes and extra latency for every request until the rebuild lands. Under replication the equivalent read just goes to another copy at full speed. Rebuild time is therefore not only a durability number; it is the length of a window in which the read path is measurably worse. ## What operators actually do about it 1. **Wait before repairing.** Most absences are transient — a restart, a brief network partition, a slow disk. Repairing immediately spends a whole blob of reads to replace a fragment that returns thirty seconds later. A delay timer before a fragment counts as permanently lost removes the bulk of wasted repair traffic. 2. **Throttle repair bandwidth.** Repair is given a bounded share of the fleet so it cannot crowd out live traffic, at the cost of a longer rebuild window. 3. **Repair by urgency, not by age.** A blob missing three of its four tolerated fragments is far closer to loss than one missing a single fragment. Prioritising by remaining redundancy puts the scarce repair budget where it changes the durability outcome. 4. **Narrow the code, or add local parity.** A smaller `k` directly cuts the amplification. Alternatively, code families add parity over small local groups of fragments as well as global parity over all of them: the common case, one missing fragment, is rebuilt from a handful of neighbours in its group, while the global parity still covers rarer multi-domain losses. The price is extra storage overhead — the saving is being spent back to buy cheap repair. ## What to watch - Do not say the repair copies the fragment from somewhere. There is nowhere to copy it from; that is the whole point of a code. - Do not think only the parity fragments take part. Any `k` survivors will do, and `k` of them are needed regardless of which fragment went missing. - Do not treat repair as a background detail. Its bandwidth is a first-class capacity input, and unthrottled it is an outage mechanism. - Do not measure durability by `n - k` alone. Time-to-repair sets the window in which the next loss matters, and degraded reads make that window visible to users.
- What makes a whole-domain failure worse than a single fragment loss?Every blob with a fragment in that domain needs its own rebuild, and each rebuild pays the full k-fold read amplification. The repair traffic is a large multiple of the lost capacity and it competes with live reads, so systems throttle repair bandwidth and rebuild the least-redundant blobs first.
- How can a code be shaped so a single-fragment repair is cheaper?By adding parity over small local groups of fragments alongside the global parity. The common case, one missing fragment, is then rebuilt from a handful of neighbours in its group rather than from k fragments, while the global parity still covers rarer multi-domain losses. The price is extra storage overhead.
saying these in an interview costs you the question
- Says the repair copies the missing fragment from another domain.
- Thinks only the parity fragments participate in reconstructing a lost fragment.
- Assumes repair traffic is negligible because only one fragment was lost.
- Treats rebuild time as irrelevant once the code tolerates enough losses.
- Believes reads fail outright while a fragment is missing.