How do you write an in-memory tier's loss window as a sentence the business signs, and what if they refuse it?
answer
- one sentence, one number, one noun
- event, window, volume, records
- quote the peak, name the owner
- the window nobody chose is the defect
- crash and planned restart differ
basics
~20 sOne sentence with a number and a noun: if this process dies we may lose up to N seconds of acknowledged writes — about M orders. If refused, the levers are flushing every write, or another home for the state.
solid answer
~50 sThe sentence has four parts: the event (this process crashes or is killed), the window (*up to N seconds of acknowledged writes*), the volume at peak (*about M records*), and what the records are in the owner's own vocabulary — orders, claims, sessions. Without the last two it is an engineering statistic nobody can sign. Name the owner explicitly and get an answer, because the defect this question is probing is that most teams never asked: whatever the store shipped with became the window by default. If the owner refuses, there are three honest responses, not one: shorten the window by flushing on every write and pay for it on the write path; move that state to a durable system of record and let the tier hold a rebuildable copy; or find that a lost record is harmless after all and sign the sentence. "We will be careful" is not on the list.
go deeper
Recall the shape of the commitment: if the process dies, up to some number of seconds of acknowledged writes is gone, and that number comes from configuration rather than from the store's reputation.
Explain how the posture and its interval produce the number, and why multiplying by the peak write rate is what makes it judgeable by someone outside engineering.
Show that you would name the owner and get a decision, and that a refusal is answered with a lever and its price — a write-path cost or another home for the state.
Price the decision across the fleet and the release cadence, separate a crash from a planned restart, and record the agreed sentence somewhere it will be found later.
## Why a sentence, and not a setting An in-memory tier's durability configuration is a business commitment written in engineering notation. Somewhere in a configuration file, a decision was made about how much of what customers were told had succeeded may vanish — and in most organisations nobody outside the team that installed the store has ever seen that decision, because it was never translated. Translating it is the work here. The translated form is one sentence: > **If this process dies, we may lose up to one second of acknowledged writes — at peak, about twelve thousand order confirmations.** ## The four parts, and why each is load-bearing 1. **The event.** "If this process dies" — a crash or kill of this single process. Being precise keeps the sentence from absorbing failures it does not describe: an acknowledged write lost because another node was promoted without it has a different cause, a different remedy and a different owner, and folding it in makes the number meaningless. 2. **The window.** *Up to N seconds of acknowledged writes* — derived from the posture in force and the interval that governs it, never from reputation or a vendor page. An **acknowledged write** is one a caller was told had succeeded; that is exactly why the sentence stings. 3. **The volume.** Seconds are not a currency outside engineering. Multiply by the acknowledged write rate at the busiest minute. 4. **The noun.** "Twelve thousand records" is still engineering; "twelve thousand order confirmations" is a thing someone can refuse. The owner's vocabulary, not the schema's. ## Who signs, and why the question exists Name the person. It is whoever answers for the consequence — the owner of the workflow the records belong to, not the platform team, and not the person who installed the store. The reason an interviewer asks is that the common state of the world is nobody having chosen. A store was installed, whatever it shipped with stayed in force, and the loss window is therefore an accident that has never been examined. Treat the unexamined setting as the defect in its own right, independently of whether the number turns out to be right: a window nobody chose is a window nobody can defend when it is spent. ## What to do with "no, that is not acceptable" There are three honest answers and one dishonest one. | Response | What it changes | What it costs | |---|---|---| | Flush on every write | the smallest window the store offers | a flush to disk on the write path for every operation — the latency the tier exists to avoid; and some stores do not offer it | | Move the state to a durable system of record first | removes this tier from the durability story for that state | an extra write on the request path, and the tier becomes a rebuildable copy | | Narrow the scope of the promise | the sentence applies to less | only works if the store can separate the state, which most cannot per key | | "We will be careful" | nothing | the window is spent later, with no record that anyone accepted it | The third row is worth dwelling on. The natural instinct is to give the important keys a smaller window and leave the rest alone, but the posture and the flush policy are generally properties of the instance rather than of a key — stores differ here, and a few let an individual call ask for stronger durability, but you should check rather than design on the assumption. If the store cannot separate them, the practical version of "narrow the scope" is to run the state that needs a different promise somewhere else entirely. ## The cost of the decision across a fleet and a release cadence A principal-level answer prices the decision beyond one instance. - **Multiply by instances.** If the answer is "flush on every write", it is paid by every operation on every instance of this tier, forever, including the ones holding state where nobody would have asked for it. - **Separate crash from planned restart.** The sentence is about a crash. A planned restart is a different event, and stores differ in what they do on a graceful shutdown — some flush and take a copy first, making the planned window effectively zero, others do not. Knowing which you have decides whether a release is a routine event or one that spends the budget. - **Count the events.** A release cadence that restarts the fleet every week turns a rare accident into a scheduled occurrence. If planned restarts are graceful and the store flushes on shutdown, that is fine; if not, the team is quietly spending the window dozens of times a year and calling it a deploy. - **Write it down where it will be found.** A sentence agreed in a meeting and never recorded is indistinguishable, a year later, from the default nobody chose. ## What the sentence must never claim Even at its smallest, this window does not make the tier a durable system of record. Persistence here bounds a restart — whether the tier comes back empty or nearly current, and how much recovery time replay at start costs. The tier remains a component built around volatile memory with a single disk beneath it. The strongest version of the answer ends exactly there: the sentence is how much bounded, survivable loss the business has agreed to, which is why this tier is never the only copy of anything that matters.
- Why not simply give the important keys a smaller loss window and leave the rest alone?Because the posture and flush policy are generally properties of the instance, not of a key — stores vary, and a few offer stronger durability per call, but most do not. Where the store cannot separate them, the workable version is to run that state on a separate deployment or in a durable system of record, rather than pretending one instance carries two promises.
- The team's releases restart every instance weekly. Does the loss window apply to those restarts?It depends on what the store does on a graceful shutdown, and stores differ: some flush and take a copy before exiting, which makes a planned restart cost effectively nothing, others do not. If yours does not, a weekly release spends the window fifty times a year, which turns a rare accident into a scheduled one nobody has priced.
- How is this sentence different from a promise about how long the tier is unavailable?They are separate currencies. The loss window is how much is gone; recovery time is how long replay at start keeps the process from serving, and an empty restart also transfers load to the system behind. Shrinking one usually worsens the other, so a commitment should state both rather than letting one stand in for the other.
saying these in an interview costs you the question
- Presents the configuration itself instead of a sentence with a number.
- Quotes seconds without converting them into records at peak.
- Treats the window as a fact of the store rather than a choice.
- Treats the write log as making the tier a system of record.
- Assumes a per-key loss window is available on any store.
- Answers a refusal with reassurance instead of a lever and its price.