skip to content

A shared tier answers 90% of 8,000 reads per second; the system of record is sized for 900. What does derivability promise here?

level: middleimportance: must knowfreq 66%

answer

  1. per value, not per second
  2. compute the steady-state source load
  3. nine times its sizing
  4. load-bearing for availability, not latency
  5. pick a position out loud

basics

~20 s

Derivability promises each value can be produced again, not that the source can serve every miss at the tier's rate. Here an empty tier sends 8,000 reads per second at a source sized for 900: an outage, not a slowdown.

solid answer

~40 s

It promises reconstructibility per value, and says nothing about rate. Steady state, the system of record is taking about 800 reads per second and has roughly 100 to spare. With the tier empty it is asked for 8,000 — an order of magnitude beyond its sizing — so requests queue, connections are held, timeouts fire, and the failure propagates to callers that were never reading the source directly. The honest reading is that this tier is load-bearing for availability, not merely for latency. Two positions are defensible: size or shed so the system of record can carry full load, or accept that the tier's availability is the service's availability and put that in the design where people can see it. What is not defensible is claiming the first while operating the second.

go deeper

for a junior

Do the arithmetic before the argument: at a 90% hit ratio the source is taking a tenth of the traffic, so an empty tier multiplies its load by ten. Compare that number against what the source is sized for.

for a middle

Explain why an overloaded source fails rather than slowing down gracefully — queueing, held connections, timeouts crossing, retries adding load — and say what the calling service actually returns to users while that is happening.

for a senior

Name the position you would take and its price: capacity at peak read rate, a defined degraded response, or an explicit admission that this tier is a critical dependency. Bring the evidence you would collect before choosing.

for a principal

The real deliverable is that the organisation stops describing a load-bearing dependency as an optimisation. Decide who owns that statement, how often the source's ability to carry full load is actually proven, and what it is worth paying to keep it true.

## Two different promises, routinely confused There are two claims people make about a derived copy, and only one of them is usually true. - **Per value:** each entry can be produced again from a system of record. This is the derivability precondition, and in this design it presumably holds. - **Per second:** the system of record can produce all of them, at the rate they are asked for, for as long as the tier is empty. This is a capacity claim, and nothing about the first claim implies it. The second claim is what "we can always fall back to the source" quietly assumes, and it is the one this design fails. ## The arithmetic, spelled out At a 90% hit ratio on 8,000 reads per second: - The tier answers about **7,200 reads per second**. - The system of record already takes about **800 reads per second** in steady state, against a sizing of 900 — roughly **100 per second of headroom**, a little over 12% of its current load. - With the tier empty, every read reaches the system of record: **8,000 per second, close to nine times its sizing**. Overload of that magnitude does not produce a proportionally slower system. It produces queueing, exhausted connections, and response times that cross the callers' timeouts. Callers time out, some retry, and the retries add load to a source already past its limit. The service is not slow; it is down, and it will stay down until the offered load drops or something sheds it. ## What this tells you about the tier A tier whose absence takes the service down is **load-bearing for availability**, not for latency. That is a legitimate thing to have, and it is a different thing from what most design documents claim. Once it is true: - The tier's availability multiplies into the service's availability, so its own failure modes become the service's failure modes. - Maintenance on the tier is user-visible by default, not a quiet task. - The dependency deserves the same review as any other component the service cannot serve without. The unacceptable state is holding both beliefs at once: describing the tier as a pure optimisation in the architecture document while operating a system that cannot serve a single second without it. ## The positions you can actually defend | Position | What it requires | What it costs | |---|---|---| | Keep the system of record able to carry full load | Capacity at or near peak read rate, and regular proof that it still can | Paying for headroom that is idle almost always | | Shed or degrade when the tier is empty | A defined reduced response, and a decision about which callers lose first | Someone has to choose, in advance, what the degraded product looks like | | Accept the tier as load-bearing | Saying so explicitly, and treating its availability as the service's | The tier now needs the availability engineering of a critical dependency | All three are real answers. What separates a strong candidate is picking one out loud, with the consequence attached, rather than asserting that the source is "there if we need it". ## What the caller experiences, in order 1. Reads that the tier would have answered reach the system of record instead, so its load rises roughly tenfold within one request cycle. 2. Response times at the source climb as work queues; the callers' own timeouts start firing. 3. Callers holding connections while they wait exhaust their pools, which makes even requests unrelated to this data slow. 4. If timed-out work is retried, the offered load rises further, and the source does not recover on its own. 5. Recovery normally requires the offered load to be reduced — by shedding, by a smaller degraded response, or by traffic genuinely falling away. ## Scope: what this question is not about This is a capacity question about the rebuild path, so two neighbouring subjects stay out of it. How you keep a derived copy in step with its source — and the techniques for damping many simultaneous misses on one value — are cache design, a separate body of practice. What a store keeps across a restart is a separate subject again. The question here is simpler and prior to all of them: when nothing is in the tier, can the rest of the system serve the traffic, and if not, has anyone written that down?

  • What measurement would you take before deciding which position to adopt?
    The read rate reaching the tier, the hit ratio, and the rate the system of record can sustain under a controlled test rather than its nameplate figure. Those three give the multiplier and the gap. Without them the discussion is instinct, and instinct consistently overestimates what a source can absorb.
  • Does a higher hit ratio make this design safer or more dangerous?
    More dangerous, other things being equal. A higher hit ratio means the system of record carries less traffic in steady state, so it tends to be sized smaller and the multiplier on an empty tier is larger. The tier becomes both more valuable and more load-bearing at once.
  • Would a per-instance copy inside each application process change this arithmetic?
    It changes where the reads come from, not whether the source can carry them. It also changes the failure shape: rather than one shared component emptying at once, each instance starts empty on deployment, so the load arrives in slices as instances restart. The capacity question survives either arrangement.

saying these in an interview costs you the question

  • Says the source is always there as a fallback, without checking its rate.
  • Confuses a value being reconstructible with the source being able to carry the load.
  • Expects overload to make responses proportionally slower rather than failing.
  • Assumes traffic will conveniently arrive slowly enough for the tier to refill.
  • Calls the tier a pure optimisation while the service cannot serve without it.