Erasure coding a blob as ten data plus four parity fragments instead of keeping three replicas — what does that change about storage cost and fault tolerance?
answer
- count stored bytes per blob byte
- fourteen tenths, not fourteen
- any four, not the parity four
- the read now touches ten places
- independence comes from placement
basics
~20 sOverhead falls from three times the blob to about 1.4 times, and tolerated simultaneous losses rise from two copies to any four fragments. The price is paid on the read and repair paths, and only if placement keeps the fourteen fragments genuinely independent.
solid answer
~50 sThree-way replication stores 3.0 times the blob and survives the loss of any two copies. A ten-of-fourteen code stores 14/10 = 1.4 times the blob and survives the loss of any four fragments, because any ten of the fourteen reconstruct the original. So it is strictly better on both headline numbers at less than half the storage. What changes underneath is the access pattern. A read now contacts ten domains instead of one, so its latency is the slowest of ten responses and its request count is ten times higher. A write computes four parity fragments and must land fragments in enough distinct failure domains that no correlated event holds more than four. And repairing a single lost fragment reads ten fragments rather than copying one. Erasure coding trades storage for network, fan-out and placement discipline.
go deeper
Hold on to the shape: replication keeps whole copies, while erasure coding keeps computed pieces where any sufficient subset rebuilds the original at a fraction of the storage.
Do the arithmetic out loud — fourteen fragments of a tenth each is 1.4 times the blob against 3.0 — and state that any four fragments may be lost, parity or data alike.
Bring the access paths and the placement rule: reads fan out to ten domains, repair reads a whole blob to restore a tenth, and no correlated failure unit may hold more than four fragments.
The call to own is where the boundary between the tiers sits: which objects are large, cold and stable enough to code, and how much read-tail and repair load the saving is worth across the fleet.
## What the two schemes actually cost | | Three-way replication | Ten-of-fourteen erasure code | |---|---|---| | Stored bytes per blob byte | 3.0 | 1.4 | | Simultaneous losses tolerated | any 2 copies | any 4 fragments | | Domains touched by a normal read | 1 | 10 | | Bytes read to repair one lost unit | 1 blob's worth per blob | 1 blob's worth per blob | | Bytes restored by that repair | 1 blob | one tenth of a blob | | Partial update of the object | write one region, three times | recompute all four parity fragments | The first two rows are why the migration happens at all: **more fault tolerance for less than half the storage**. Four tolerated losses beat two, and 1.4 times beats 3.0 times. At any real scale that difference is the budget line that funds the project. ## Where the durability actually comes from The code's guarantee is combinatorial: any `k` of the `n` fragments determine the blob, so the object survives any `n - k` absences. That statement is unconditional about *which* four go — it does not matter whether they are parity fragments or data fragments, a point worth being explicit about because the asymmetry of the names suggests otherwise. What it is not unconditional about is **independence**. The arithmetic assumes the four losses are four separate events. If five fragments share a power feed, a firmware version, a rack, or a deployment that rolls forward together, one correlated event ends the blob while the code's arithmetic still says it should have survived. So the real requirement is a placement rule: *no single correlated failure unit may hold more than `n - k` fragments of the same blob*. Replication has the same exposure, but with only three copies to place, people usually get it right by accident; with fourteen fragments the constraint has to be designed and enforced. ## What it costs on the access paths - **Reads fan out.** Reconstructing any byte range needs `k` fragments, so an ordinary read contacts ten domains. Its latency is the slowest of ten responses, not the median, and the request rate against the fleet rises tenfold. Systems soften this by requesting a few extra fragments and decoding from whichever `k` arrive first, spending bandwidth to cut the tail. - **Writes compute and scatter.** The write path computes four parity fragments and must place fourteen pieces in fourteen acceptable locations. Acknowledging a write early means acknowledging before the full fragment set is durable, which is a real availability-versus-durability decision, not a detail. - **Repair reads wide.** Restoring one lost fragment requires fetching `k` surviving fragments — a whole blob's worth of network traffic to regenerate a tenth of a blob. Replication restores a lost copy by copying one copy. - **Updates are expensive.** Every parity fragment is a function of all ten data fragments, so changing one byte means recomputing and rewriting all four parity fragments. Coded stores therefore prefer write-once or append-only objects and re-encode whole blobs rather than patching them. ## Where replication still wins Erasure coding is not a strict upgrade, and a good answer says where it loses: 1. **Small objects.** Splitting a tiny object into ten fragments produces ten pieces each smaller than the per-fragment metadata and minimum I/O unit. The overhead ratio inverts and the object costs more coded than replicated. 2. **Hot objects.** A read served from a single replica beats a ten-way fan-out on both latency and request cost, and the difference shows up directly in the tail. 3. **Frequently mutated objects.** The parity-rewrite cost per update makes a coded representation a poor fit for anything edited in place. The common shape is a tiered one: replicate on write, then re-encode into the coded tier once the object is cold, large enough, and unlikely to be modified. ## What to watch - Do not quote 14 times or 5 times as the overhead. Fragments are a tenth of the blob each; fourteen of them is 1.4 blobs. - Do not claim the parity fragments are the only ones that may be lost. Any four fragments may go. - Do not present the durability gain without the placement rule. Independence is delivered by placement, not by the code. - Do not migrate the whole corpus. Small, hot and mutable objects belong on replication.
- Does four tolerated losses always beat three replicas on real durability?Only if placement makes those losses independent. Four beats two arithmetically, but if a correlated event — shared power, one firmware version, a rollout that moves together — takes five domains at once, the coded blob is gone while a differently spread replica set might survive. The code assumes an independence that placement has to deliver.
- What does erasure coding do to the read path?An ordinary read contacts k domains rather than one, so its latency becomes the slowest of k responses and its request count multiplies by k. Systems mitigate this by requesting a few extra fragments and decoding from whichever k return first, trading bandwidth for a shorter tail.
- Why is updating part of an erasure-coded blob expensive?Each parity fragment is a function of all k data fragments, so touching one data fragment forces recomputing and rewriting every parity fragment. Coded stores therefore favour write-once or append-only objects, and re-encode a whole blob rather than patching it in place.
Three-way replication is three full photocopies of a page. A ten-of-fourteen code cuts the page into ten strips and computes four more, and any ten strips at all rebuild the page — fourteen strips of a tenth of a page is 1.4 pages of paper, not fourteen.
saying these in an interview costs you the question
- Says fourteen fragments means fourteen times the storage.
- Believes only the four parity fragments may be lost safely.
- Presents erasure coding as strictly better with no access-path cost.
- Ignores placement and assumes fragment losses are independent by default.
- Recommends coding small or frequently read objects alongside bulk data.