skip to content

Why does storage attached in one zone pin a data-owning copy to that failure domain, and what does losing it cost?

level: seniorimportance: should knowfreq 47%

answer

  1. the data decides where the copy runs
  2. attachment narrows the candidate hosts
  3. domain gone, nothing left to place it on
  4. the process starts fast, the data moves slow
  5. an independent copy elsewhere removes the pin

basics

~20 s

Placement follows the data: a copy can only run where its backing store can be attached, so the candidate hosts shrink to one zone. If that zone is lost the copy has nowhere to go, and recovery is bounded by moving or rebuilding the data, not by starting the process.

solid answer

~50 s

An interchangeable copy can be placed on any host with capacity, which is why a lost host is a non-event. A copy that owns data adds a constraint the scheduler cannot relax: the host it lands on must be able to attach that copy's backing store. Block storage of this kind is typically attachable only within the zone that holds it, so the candidate set collapses from the whole cluster to one failure domain - that is data gravity, and it applies to placement long before it applies to migrations. When the zone goes, the copy is unplaceable: the platform keeps trying and keeps failing, because the one thing it needs is not in any remaining domain. The cost is not the restart. It is transferring or rebuilding the data elsewhere, which is measured in the size of the data over the rate you can move it.

go deeper

for a junior

Recall that a workload holding data can only run where that data is attached, so it is not free to move around the cluster the way a stateless copy is.

for a middle

Explain placement as a filter and storage attachment as one of its constraints, and say how block storage, shared file storage and a host-local path each narrow it differently.

for a senior

Quantify the recovery: name the data size, the transfer rate, and the copy elsewhere that bounds the outage, rather than quoting a process start-up time.

for a principal

Decide which workloads are allowed to be single-domain, price the second stored copy against the downtime it prevents, and make sure that choice is recorded before an event, not during one.

## Placement is a filter, and storage is one of the filters When a scheduler places a workload, it eliminates hosts that cannot satisfy the copy's requirements and picks among what is left. For an interchangeable copy the filters are cheap - enough free capacity, the right kind of host, no rule forbidding it - and there are usually many survivors. A data-owning copy adds one more: **the host must be able to attach this copy's backing store**. How tight that filter is depends on the kind of store: - **Block storage attached to one host at a time** is typically reachable only inside the zone that holds it. The filter reduces the candidate set to hosts in that zone, and the copy is effectively resident there. - **Shared file storage reachable from several hosts** widens the filter to wherever that storage is exported. The copy can move further, but the data is now behind a network path, with the latency and the availability of that path added to every read and write. - **Storage that exists only on the host** - a path on that machine's own disk - is the tightest filter of all: the copy can run on exactly one host, and a host failure is a data event. This is data gravity in its everyday form. The mass is not a migration project; it is the fact that the running copy can only be where its data is. ## What losing the domain actually costs | | Interchangeable copy | Data-owning copy in a lost zone | |---|---|---| | Placement after the loss | any surviving host with capacity | none, until data is available elsewhere | | What the platform does | starts a replacement, seconds | retries placement, fails, retries | | Recovery bounded by | process start-up time | moving or rebuilding the data | | Human decision needed | none | which copy of the data, restored where | The important line is the third. Starting a process takes seconds whether the data is 2 GB or 2 TB. Restoring 400 GB over a path that sustains 100 MB per second takes over an hour before anything is serving, and that assumes the copy is current, reachable and already outside the lost domain. Any recovery estimate that quotes a start-up time for a data-owning workload is measuring the wrong thing. The second line is worth saying precisely too, because it is often misread as a platform fault: the platform is behaving correctly. It keeps attempting placement, and every attempt fails the storage filter, because no surviving host can attach a store that is gone. There is nothing for a control loop to fix; the missing input is the data. ## Loosening the pin, and what each option costs 1. **Keep an independently stored, current copy of the data in another domain.** This is the only option that genuinely removes the pin, because it gives the copy somewhere to be. It costs a second set of storage, the work of keeping the second copy current, and the honest admission that switching to it is a decision with a data-loss number attached - how far behind the second copy was when the first one went. 2. **Use storage reachable from more than one domain.** The copy can then be placed wider. You pay in latency on every operation, and you have replaced one failure domain with another - the shared storage service itself, whose loss takes every copy that depends on it rather than one zone's worth. 3. **Accept the pin and make the domain the unit of failure.** Entirely legitimate for workloads whose tolerance for being down matches how fast the domain typically comes back. The mistake is not choosing this; it is choosing it silently, so the first zone event is also the first time anyone discovers the workload was single-domain. ## The reasoning to bring to an interview The useful chain is short and it should be said in this order: storage constrains placement, placement constrains where a copy can recover, and the recovery clock is set by the data size and the path back rather than by the workload spec. Then name the trade-off you actually made, and the number that goes with it. The weak answer treats a lost zone as a scheduling problem that the platform will resolve on its own; the strong one names the data as the thing that was lost and the copy elsewhere as the only thing that shortens the outage.

  • Does shared file storage reachable from several hosts remove the failure-domain problem?
    It moves it. The copy can be placed wider, which is a real gain, but every read and write now crosses a network path with its own latency, and the shared storage service becomes a failure domain of its own - one whose loss takes every workload depending on it rather than one zone's worth. It is a different trade, not an escape.
  • A data-owning copy has been pending placement since a zone was lost. What is worth checking, and what is not?
    Worth checking: whether a usable copy of that member's data exists outside the lost domain, and how long restoring it takes. Not worth checking: scheduler tuning, capacity in the surviving zones or placement rules. The filter that fails is storage attachment, and no amount of free capacity elsewhere satisfies it.

Moving a desk to another floor takes one person and a lift; moving the archive room the desk works from is a project. In practice the archive decides which floor the desk lives on.

saying these in an interview costs you the question

  • Says a data-owning copy can be placed on any host with capacity
  • Treats a lost zone as a scheduling problem the platform will fix
  • Estimates recovery from how fast the process starts
  • Believes shared file storage removes the trade-off entirely
  • Counts a second copy on the same device as redundancy
  • Discovers the workload was single-domain during the outage