skip to content

When is the right answer to accept an empty restart and make restarting cheap, rather than engineering restores and pre-loading for a whole fleet?

level: principalimportance: nice to knowfreq 34%

answer

  1. three places to put the cost
  2. measure the headroom, do not assume it
  3. nodes times cadence, not one event
  4. expensive to rebuild changes the answer
  5. write the sentence with two numbers

basics

~20 s

Accept it when the tier holds only reconstructible entries and the system of record behind it has been measured absorbing the full read rate. Engineer around it when that capacity is a hard ceiling, or entries have no other home.

solid answer

~40 s

The cost of a restart has to land somewhere: permanent headroom on the system of record behind the tier, effort at restart time, or an accepted degradation window somebody has signed. Accepting an empty restart is right when the entries are all reconstructible, the system behind has been measured carrying the full arrival rate for as long as a fill takes, and restarts are frequent enough that a per-restart procedure becomes a permanent tax. It stops being right when the capacity behind the tier is a hard ceiling, when entries are expensive to rebuild rather than merely to fetch, when fleet size and release cadence multiply the empty restarts per month, or when the tier holds state with no other home - then a cheap restart is data loss, not a latency event.

go deeper

for a junior

Know that some teams deliberately let the tier come back empty, and that this only works when everything it held can be rebuilt from somewhere else.

for a middle

Explain the inputs to the decision: read rate, hit ratio, fill time and the spare capacity of the system of record behind the tier.

for a senior

Show that you would measure one real restart before deciding, and that you would multiply the event by the fleet and the release cadence rather than reasoning about a single restart.

for a principal

Own the choice explicitly - permanent headroom, restart-time engineering, or a written degradation window - and rule the tier out of the durability story for state that cannot tolerate any window.

## Where the cost can go A restart has a price and there are only three places to put it: 1. **Permanent headroom behind the tier** - size the system of record so the empty-restart peak is simply absorbed. You pay every day for an event that happens occasionally. 2. **Effort at restart time** - restoring from a copy, pre-loading a chosen key set, staggering the cycle by order and pacing. You pay in engineering, in procedure, and in how long a restart takes. 3. **An accepted degradation window** - the tier comes back empty, things are slower or partly unavailable for a stated period, and somebody has agreed to that in advance. The defect is not choosing the wrong one. The defect is not choosing, which in practice means option three without the agreement. ## When accepting the empty restart is right - **Everything the tier holds is reconstructible.** Nothing is lost by an empty restart except time; the entries are copies of something durable. - **The system behind has been measured absorbing the full arrival rate**, for at least as long as a fill takes. Measured, not assumed - the multiplier is one over the miss rate, and a tier with a high hit ratio is hiding a large multiple. - **Restarts are frequent.** If the tier is restarted by every deployment, a per-restart procedure is not an occasional cost, it is a permanent one, and making the restart itself cheap beats making it elaborate. - **A restore would take longer than the shock lasts.** On a large tier, reading back a copy or replaying a **write log** can leave the node unavailable longer than an empty node would have been merely degraded. Then keeping the copy buys nothing at restart time. - **The tier is growing.** Recovery time scales with how much is stored, so a procedure that works today gets worse exactly as the system succeeds. ## When it is not right - **The capacity behind the tier is a hard ceiling.** A licensed engine, a rate-limited third party, or a shared engine other workloads depend on cannot be told to absorb twenty times its read load for ten minutes, no matter how briefly. - **Entries are expensive to produce, not merely to fetch.** If an entry is the result of a computation, a join across services or a paid call, the fill is not a read spike - it is a bill and a latency cliff. - **The fleet multiplies the event.** Empty restarts per month is roughly *nodes x releases per month*, plus involuntary restarts. A modest fleet on a weekly cadence turns "the restart is fine" into a recurring capacity line item. - **The state has no other home.** Then an empty restart is data loss, and "cheap restart" is a decision about losing data, not about latency. The honest options are a **posture** that keeps something, or moving that state off a volatile tier entirely. ## The arithmetic that decides it | input | how you get it | why it matters | |---|---|---| | arrival rate at the tier | measured at the tier | sets the absolute spike | | hit ratio | measured at the tier | the multiplier is one over the miss rate | | fill time | measured on one real restart | how long the spike lasts | | headroom behind the tier | measured, at the trough and the peak | whether it fits at all | | restarts per month | nodes x cadence, plus involuntary | turns one event into a rate | If the spike fits in the headroom at the busiest hour you would ever restart in, accepting the empty restart is defensible and you should take the simpler system. If it fits only at the trough, you have a scheduling constraint, not a free lunch. If it does not fit at all, the decision is already made for you. ## The sentence you have to be able to write The durability discussion has a sentence with a number in it - we may lose up to so many seconds of acknowledged writes. The restart discussion needs its own: *after any single node restart, the system of record behind the tier absorbs up to N times its normal read rate for up to M minutes, and we have verified that it can.* If nobody can fill in N and M, the team did not choose to accept an empty restart; it defaulted into one and will find out the numbers during an incident. ## What varies, and what it means for the choice Some stores in this class keep nothing across a restart by design and offer no setting that changes it. On those, "accept an empty restart" is not a decision you made - it is a property of the store you picked, and the real decision is whether that store suits this data. Others offer a point-in-time copy, a write log, or both, and there the choice is live. Either way the conclusion is the same one this whole category keeps arriving at: persistence here bounds how long a restart hurts; it does not turn the tier into the place the only copy of something lives.

  • Why does fleet size change this decision rather than just repeating it?
    Because the event rate is what the system behind actually experiences. Empty restarts per month is roughly nodes multiplied by release cadence, plus involuntary restarts, so a design that shrugs off one restart a quarter may be absorbing several a week. At that rate the elevated load stops being an exception and becomes part of the normal capacity model.
  • What changes when entries are expensive to compute rather than expensive to fetch?
    The fill stops being a read spike and becomes a cost and latency event. Rebuilding a computed or purchased entry cannot be absorbed by adding read capacity behind the tier, so pre-loading a chosen key set ahead of traffic, or a posture that keeps something, moves from nice to necessary.
  • Is there a case where keeping a copy makes the restart worse?
    Yes. On a large tier, restoring can leave the process unavailable for longer than an empty node would have been merely degraded, and entries come back stale. If the data is reconstructible and the system behind can carry the fill, coming back empty and serving immediately is the better outcome.

saying these in an interview costs you the question

  • Assumes headroom behind the tier without ever measuring it.
  • Treats one restart as the event, ignoring fleet size and release cadence.
  • Says a kept copy is always better than an empty restart.
  • Proposes accepting an empty restart for state with no other home.
  • Cannot state the degradation window with numbers anyone agreed to.
  • Argues persistence here makes the tier a safe system of record.