An object store advertises durability with many nines, but ingest requests fail for an hour - which promise did that figure never make?
answer
- surviving against reachable
- two promises, two units
- an outage attacks the path
- many nines is a loss probability
- buffer and retry answer the outage
basics
~20 sDurability is about bytes surviving; availability is about bytes being reachable now. A many-nines durability figure estimates how unlikely it is that the store loses an object, and says nothing about an hour of failed requests.
solid answer
~50 sTwo different promises are in play. **Durability** answers whether the bytes will still be there later: it is the store's estimate of the chance that an object it already acknowledged can no longer be reconstructed, and it is bought with copies and with repair of copies that are lost. **Availability** answers whether you can read or write right now: it is the fraction of time the request path answers successfully, and providers publish it as a far weaker number, an availability commitment in the high nines at best. An hour of failed uploads attacks the second and leaves the first alone - the objects already stored are untouched by the outage, they are simply not reachable. So the durability figure tells me nothing useful here; what matters is whether my ingest process buffers and retries, so the hour becomes a stall rather than lost events.
go deeper
Be able to say the two apart in one line: durability is bytes surviving, availability is bytes reachable now. Quoting the many-nines figure when asked about an outage is the stumble interviewers are listening for.
Explain the units. Durability is a probability of losing an acknowledged object over a period; availability is the fraction of time the request path answers. Then say why an outage moves the second and not the first.
Show what you do about each. Durability is bought with the spread of copies; availability is absorbed by retries, buffering and a decision about what the caller does while it waits. Name which one your ingest path actually needs.
Frame it as spend. Reachability costs continuously and mostly protects revenue in the moment; byte survival is a copy decision with a smaller ongoing bill. Say which failure the organisation can tolerate and let that set both budgets.
## Two promises that get quoted as one An object store makes two separate promises about what you put in it, and they are measured in different units. **Durability** is about the bytes surviving. It is the store's estimate of the probability that an object it has acknowledged can still be reconstructed later. It is defended by keeping more than one copy of the object and by repairing or re-creating copies when the media or the machine under them is lost. Providers publish it as a figure with many nines, stated per object over a period. **Availability** is about the bytes being reachable. It is the fraction of time the store's request path answers your reads and writes successfully. It is defended by redundancy in front of the data: request handlers, name resolution, the network between you and the store, and whatever routes a key to a copy. The published number is always much weaker than the durability one. | Axis | Durability | Availability | | --- | --- | --- | | The question it answers | Will the object still exist later? | Can I read or write it at this moment? | | Unit | Probability of losing an object over a period | Fraction of time requests succeed | | Defended by | The number and spread of copies, plus repair | Redundancy of the request path, retries, a second path | | Broken when | No copy of the object can be reconstructed | Requests error or time out although the bytes are intact | | Your lever | The replication scope you choose | What your client does while the endpoint refuses | ## Why an outage does not move the durability number During an incident that makes the endpoint refuse requests, very little is happening to your objects. The media still holds them, the copies still exist, the repair process still runs. What failed is the path: the front end that terminates your connection, the component that resolves a key to a copy, or the plumbing between you and them. When service returns, your reads return the same bytes as before. That is why answering an outage question with a durability figure is the classic stumble - the two failures have almost nothing in common, and a store can be strong on one and weak on the other at the same time. The converse is worth saying out loud too. Being reachable says nothing about surviving: a store that answers every request promptly may still be keeping a single copy of each object in one failure domain. ## What this means for a raw-event landing zone Take the common shape where an ingest process writes raw events into an object store all day and nightly jobs read them back. An hour of failed writes is an availability event, and every decision it forces sits on the writing side: - **What happens to events the process has already accepted.** They are in your memory or on your local disk, not in the store, so no store promise covers them at all. Buffering them, or leaving them unacknowledged in the upstream queue, is what turns the hour into a stall instead of a hole in the day. - **How long you can stall.** The backlog you can absorb is bounded by your buffer and by the upstream retention, not by the store's nines. - **Whether the nightly job still runs.** If reachability returns before the job starts, the outage costs nothing downstream. None of those is informed by the durability figure. The reverse case is a genuine durability claim: if a nightly job cannot find an object that ingest was told had been written, that is worth investigating as loss - starting with whether the write was really acknowledged, because an unacknowledged write is outside the promise. ## Where confusing the two costs you - Reporting an outage to stakeholders as data loss when nothing was lost. - Starting a recovery that was never needed because reads were failing. - Sizing no client-side buffer, on the grounds that the store is described as extremely reliable. - Assuming a many-nines durability figure implies copies in more than one place, when the spread of copies is a separate setting with its own price. ## How to answer it out loud 1. Name both properties and their units in one sentence each. 2. Say which one the scenario attacks: failed requests are an availability event, and the stored bytes are not the thing at risk. 3. Put the remedy where it belongs - in the client that buffers and retries, not in a storage setting - and add the honest caveat that neither promise defends against a bad write of your own, which is a third thing again.
- The store is unreachable for an hour. What should the ingest process do with events it has already taken in?Hold them rather than discard them: keep them buffered locally or leave them unacknowledged upstream, and retry with backoff. Those events sit outside every store promise until a write is acknowledged, so the only thing protecting them is your own retention. The honest design question is how long the buffer lasts, because that is the real outage budget.
- Does a higher durability figure make a store more available?Not by itself. Copies spread across failure domains can help both properties, since a request can be served from a surviving copy, but the request path in front of the data is what usually fails in an outage, and more copies do nothing for that. Treat them as separate levers that happen to overlap.
- Why is the published availability number always far weaker than the durability number?Because availability covers the whole live request path - front ends, routing, name resolution, network - and every one of those can fail briefly without endangering a single byte. Durability only has to be true over a long period and is repaired continuously in the background, so it can be quoted with many more nines.
A warehouse can promise it will never lose your crates and still have its doors shut all afternoon. Both statements are true at once, and only one of them is about the crates.
saying these in an interview costs you the question
- Says a many-nines durability figure means the store is almost never down
- Treats the durability figure and the availability commitment as the same number
- Assumes failed uploads mean data already stored was lost
- Claims high durability removes the need for the client to retry
- Answers an outage question by quoting the durability nines