A team wants to hold state with no other home on an in-memory tier whose loss window is one second — how do you test whether that is acceptable?
answer
- seconds mean nothing; records mean something
- use the busiest minute, not the average
- reconstructible copy or no other home
- one lost record: reversible, detectable, bounded
- no window tolerable means elsewhere
basics
~20 sConvert the window into records at peak, then ask what one lost record costs and who finds out. For a reconstructible copy the window is a performance question; for state with no other home it is a correctness budget.
solid answer
~50 sThree steps. First turn seconds into volume: the acknowledged write rate at the busiest minute times one second is the number of writes a crash discards. Second, classify what those writes are. A **reconstructible copy** of something held in the system of record behind the tier loses nothing permanently — a crash costs an empty restart and load transferred to the system behind, which is a capacity question. **State with no other home** is different: those records are simply gone, and the window is a correctness budget. Third, name the observable consequence of one lost record — a claim granted twice, a quota reset, a request re-processed — and ask whether it is reversible and whether anyone detects it. If the honest answer is that no window is tolerable, the conclusion is not a shorter timer: that state should be written to a durable system of record first, with the tier holding a copy that can be rebuilt.
go deeper
Recall that what a crash costs depends on what the entries are: a copy of something stored elsewhere can be rebuilt, while state that exists only here is simply gone.
Explain the arithmetic — window in seconds times the acknowledged write rate — and why the peak rate, not the average, is the number worth quoting.
Demonstrate the classification and the single-record test: reversible, detectable, bounded. Say plainly when the verdict is that this state does not belong on the tier alone.
Frame it as a budget someone signs rather than a setting an engineer picks, and be willing to conclude that the honest answer is a durable system of record in front of the tier.
## The question behind the question "Is a one-second loss window acceptable?" cannot be answered from the number. One second of acknowledged writes is trivially acceptable for some state and unacceptable for others, and the work is entirely in showing which this is. The test has three steps, and a candidate who has run one of these tiers does all three without being asked. ## Step one: seconds into records A window is a span of time; what a business can judge is a count. Multiply the window by the acknowledged write rate — at the busiest minute of the week, not the daily average, because a crash is not considerate about timing. - at 200 writes a second, a one-second window is a couple of hundred records; - at 40,000 writes a second it is 40,000; - and if the write rate is bursty, the number to quote is the burst, since that is when the tier is most likely to fall over. This step alone changes conversations. "One second" sounds like nothing; "about forty thousand acknowledged writes" does not. ## Step two: classify what is lost There are two genuinely different cases and they have different tests. | What the entries are | What a crash actually costs | The question to ask | |---|---|---| | A reconstructible copy of state held in the system of record behind the tier | an empty restart and the load that transfers to the system behind while the tier refills | can the system behind absorb it? | | State with no other home — a claim, a counter, a record that something was already handled | those records are gone and cannot be re-derived | what does one missing record do, and who notices? | The premise of this question puts the workload in the second row, which is where the loss window stops being a performance concern and becomes a correctness budget. The distinction is worth stating explicitly in the interview, because a candidate who blurs it will either over-engineer a rebuildable copy or under-engineer state that has nowhere else to live. ## Step three: name the consequence of one lost record Aggregates hide the damage. Force the question down to a single entry and ask three things about it: 1. **Is it reversible?** A miscounted view is self-healing at the next write. A claim that a second worker can now also take is not. 2. **Is it detectable?** Damage nobody observes is still damage, and it is worse, because it never enters an incident review. 3. **Is it bounded?** One lost record that causes one duplicated action is a different risk from one lost record that unblocks an unbounded amount of duplicated work. At this point the answer usually stops being about the tier. If a single lost record produces an irreversible, undetectable effect on a customer, the honest verdict is that the tier's window is not the right budget for it, regardless of how small you make the number. ## What to do when no window is acceptable The reflex is to shrink the window: flush on every write rather than on a timer. That is a real lever and it has a real price — a flush on the write path for every operation, which is the latency the tier exists to avoid — and on some stores it is not offered at all. But note what it does not do: it does not make the tier a durable system of record. The disk under the process is one failure domain, and the acknowledgement-to-flush ordering varies by store even in that posture. The structural answer is to move the decision off this tier: - **write to a durable system of record first**, then let the tier hold a fast copy that can be rebuilt — which converts the state back into the first row of the table; - **make the operation tolerant of a lost record**, so that losing one degrades into a retry or a re-check rather than a wrong outcome; - **accept the window explicitly**, in a sentence somebody signs, when the cost of a lost record is genuinely small. Those are the three honest outcomes. "We made the timer smaller and stopped worrying" is not one of them. ## Keeping the test in scope Two neighbouring failures get dragged in here and should be kept out, because they have different causes and different remedies: - an acknowledged write lost because a **different node** was promoted without it — that is a property of how the tier is replicated, not of the flush policy, and shrinking the flush timer does nothing for it; - the **recovery time** and the load an empty restart transfers — a real cost of a crash, but a separate currency from writes lost. The loss-window test is about flush and copy timing on a single process, and keeping it there is part of answering well. ## What a strong answer sounds like "At peak we acknowledge about twelve thousand writes a second, so a one-second window is about twelve thousand records. They are claims with no other home, and losing one means a second worker can take the same claim — reversible, but visible to a customer. I would either write them to a durable system of record first and treat the tier as a rebuildable copy, or accept the window in writing after showing that number to the person who owns the workflow." That is a budget, a classification, and a decision — not a setting.
- The team proposes flushing on every write instead. Does that make the tier safe for state with no other home?It shrinks the window and costs a flush on the write path for every operation, which is the latency the tier exists to avoid. It does not confer a durable guarantee: the disk beneath the process is a single failure domain, and stores differ in whether the acknowledgement precedes the flush even in that posture. It is a smaller budget, not the absence of one.
- Why quote the write rate at the busiest minute rather than the daily average?Because the window is multiplied by whatever rate is in force when the process dies, and pressure-induced crashes cluster at peak. An average understates the exposure exactly where it matters, and the number shown to whoever signs the sentence should be the one they would face on the worst afternoon.
saying these in an interview costs you the question
- Judges the window from the seconds without converting to records.
- Uses the average write rate rather than the peak.
- Treats a reconstructible copy and state with no other home identically.
- Concludes that a smaller flush interval makes the tier a system of record.
- Brings failover of another node into a single-process loss-window test.
- Calls the window acceptable without naming what a lost record does.