skip to content

What extra failure does an object store survive as its copies move from one zone, to several zones, to another region?

level: middleimportance: should knowfreq 60%

answer

  1. how far apart the copies sit
  2. each step buys one failure class
  3. zones nest inside a region
  4. the far copy is asynchronous
  5. check the configured scope, not the figure

basics

~20 s

Copies inside one zone survive a failed disk or machine. Copies spread across zones in a region survive losing a whole zone. A copy in another region survives losing the region. Each step costs more and is usually a setting you choose.

solid answer

~50 s

Replication scope is the question of how far apart the copies are, and each widening buys one more class of correlated failure. Copies kept inside a single zone cover the everyday case: a disk that fails, a machine that dies, media that rots - the store re-creates the lost copy from a surviving one. Spreading copies across zones in the same region covers the loss of a whole zone, which is a genuinely correlated event that within-zone copies do not survive. A copy in a second region covers the loss of the region itself, and because the distance makes synchronous writes expensive, that copy is normally maintained asynchronously and therefore trails the newest objects. The important operational detail is that the wider scopes are typically something you configure and pay for per copy and per byte moved, so the first thing I check on an existing store is which scope it is actually running, rather than assuming.

code

yaml · 9 lines
yaml
store: raw-events-landing
replication:
  scope: acrossZonesInRegion   # copies in separate failure domains of one region
  mode: synchronous            # acknowledged only after these copies exist
additionalCopies:
  - scope: secondRegion
    mode: asynchronous         # applied after the write is acknowledged
    expectedLag: minutes
    appliesTo: newObjectsOnly  # existing objects need a backfill

go deeper

for a junior

Get the nesting right first: zones are separate failure domains inside one region, and a region is the bigger, independent unit. Copies in one zone die with that zone.

for a middle

Explain what each widening buys and what it costs, and know that the far-region copy is normally asynchronous. Say that the scope is a setting you should read off the store rather than assume.

for a senior

Demonstrate the operational checks: which scope is actually configured, whether it was backfilled over existing objects, and what the copies add to the bill for a high-volume raw zone.

for a principal

Treat it as a spend decision against regenerability. Data the upstream can replay rarely justifies a doubled storage bill; name the subset that genuinely cannot be rebuilt and replicate that.

## What a copy is for Durability is not a property of an algorithm, it is a property of how many independent things have to fail before an object cannot be rebuilt. So every durability discussion reduces to one decision: **how far apart do the copies sit**. That distance is the replication scope, and each widening covers one more class of failure that the narrower scope treats as a single event. ## The three scopes | Scope | Survives | Usual write mode | What it costs | | --- | --- | --- | --- | | Copies inside one zone | A failed disk, a failed machine, silent media decay | Synchronous, before the write is acknowledged | Usually included in the base per-gigabyte price | | Copies across zones in one region | The loss of a whole zone, including power, cooling or network for that site | Usually synchronous; the zones are close enough | A higher per-gigabyte rate, plus traffic charged between zones for some access patterns | | A copy in a second region | The loss of an entire region | Normally asynchronous, because the distance makes synchronous writes slow | Storage paid twice, plus a charge for the bytes moved between regions | Two details matter more than the table. First, **the failure domains are nested, not parallel**. Zones are separate failure domains inside one region; a region is the larger, independent unit. Copies in three zones do not help when the region is the thing that is gone, and a copy in another region is far away, which is precisely why it is expensive and why it lags. Second, **the far copy is asynchronous in almost every product**. The write is acknowledged once the near copies are durable, and the far copy is applied afterwards. That is a deliberate trade: you get the region-loss protection without paying the latency of a long-distance round trip on every write, and you accept that the far copy is behind by a variable amount. ## Automatic against configured The most common mistake on a real store is assuming the scope rather than checking it. Products differ here - some spread copies across zones in a region by default and charge for it in the base price, others keep everything in one zone unless you ask, and the far-region copy is essentially always something you turn on, name a destination for, and pay for separately. The engineering habit that matters is: 1. Read the store's configured scope, not the marketing figure. 2. Confirm whether the wide scope applies to **new objects only** or was backfilled over what is already there, since rules that start applying today do not retroactively copy yesterday. 3. Confirm what the scope covers: whole objects, or objects matching a subset such as a key prefix. ## What no scope buys you Widening the scope defends the bytes against hardware and against sites. It does nothing about a bad write, because replication copies whatever you asked for to every copy in turn. It also does not by itself make the store more reachable - copies across zones can let a request be served from a surviving zone, but the request path in front of the data is a separate piece of machinery with its own failure modes. ## Choosing for a raw-event landing zone For a landing zone written by ingest all day and read by nightly jobs, the reasoning runs like this: - The events are **reproducible or not**. If the upstream can replay a day, the value of a far copy drops sharply and you are mostly paying for convenience. - The volume is **large and cold**. Cross-region copies of a high-volume raw zone are one of the easiest ways to double a storage bill for data that is read once by a nightly job. - The read pattern is **regional**. If every consumer of the raw zone lives in one region, a second-region copy exists for the loss of the region and nothing else, and should be justified on exactly that. A reasonable default for that shape is copies across zones in the region, plus a far copy only for the subset that cannot be regenerated. Saying that trade out loud - what survives, what it costs, what is regenerable - is the answer an interviewer is listening for, rather than the reflex of replicating everything everywhere.

  • Why is a copy in a second region normally asynchronous when copies across zones can be synchronous?
    Distance. Zones sit close enough that a write can wait for all of them without a latency the caller notices, so the acknowledgement can mean every near copy exists. A second region is far enough that waiting for it would add real round-trip time to every write, so the copy is enqueued and applied afterwards, which is why it trails.
  • You enable a wider replication scope on a store that already holds a year of objects. What have you actually changed?
    Usually only the future: the setting applies to objects written from that point on, and everything already there keeps its old spread until you run a backfill. That gap is a common surprise during an audit, because the store reports the wide scope while most of its bytes do not have it.
  • Do more copies improve how fast reads are served?
    Not reliably. Copies exist to survive failures, and whether a read can be steered to the nearest one depends on the product and on where the reader sits. Treat read performance as a separate question about the access path and the tier, not as a side effect of the replication scope.

saying these in an interview costs you the question

  • Thinks copies in one zone survive losing that zone
  • Describes a region as something inside an availability zone
  • Assumes every store replicates across zones by default
  • Believes a far-region copy is kept synchronously in step
  • Expects a newly enabled scope to cover objects written last year
  • Calls cross-zone replication free because it stays in one region