A service reads a derived copy from a shared tier that stops answering; what decides whether its requests degrade or fail?
answer
- two outages wear one name
- derived copy or the only copy
- is there a route to the source
- the source takes the full rate
- up but refusing is also degraded
basics
~20 sThree facts decide it: whether the entries are a derived copy with a system of record behind them, whether the read path has a route to that source, and whether the source can absorb the tier's full read rate.
solid answer
~40 sBecause the entries are a **derived copy**, the data still exists at the system of record, so degradation is possible in principle. Whether it happens depends on the code: if the read path has a route to the system of record, requests are served more slowly but correctly; if it assumes the tier answers, the design fails with extra steps. Then the load question - every read the tier was absorbing lands on the source at once, with no ramp, and the source may have been sized after the tier existed. And *stopping answering* is not one behaviour: unreachable, slow, emptied by a restart and at its memory ceiling are four different events. Stores differ at the ceiling too - some reclaim entries, some refuse writes while still serving reads.
go deeper
Remember that the tier can stop answering at any moment and that the data behind a derived copy still exists elsewhere. The question a reviewer will ask is what your code does when the lookup returns nothing.
Explain the mechanics: the route to the system of record has to exist on that path, the full read rate arrives there at once, and the caller waits rather than failing instantly when a component is unreachable.
Show production judgement by naming the rebuild-path load and the mode that actually hurts - a tier that is slow rather than down - and by refusing to treat one store's ceiling behaviour as the model.
Treat the blast radius as the real subject: one shared tier means many services degrade in the same minute and many sources take their unfiltered rate together, which is a different incident from one service having a bad day.
## Start with what the tier holds The scenario says the entries are **derived copies**: faster copies of data a **system of record** still holds authoritatively. That one fact is what makes degradation possible at all. Had the entries been the **sole home for ephemeral state** - facts with no other holder - the same outage would not be a latency event but a loss of facts, and no read path could recover them. Every answer about a tier that has stopped answering begins by saying which of the two it is, because the correct answers are opposites. ## What the caller experiences, in order 1. A request that would have consulted the tier does not get an answer promptly. An unreachable component typically does not refuse politely; it leaves the caller waiting, holding a thread and a connection, until the caller stops waiting. How that waiting is bounded is a separate subject, but the waiting itself sits on the hot path of every request that consults the tier. 2. If the code has a route to the system of record from that point, the read is served - slower, but correct. The user may see a slower page rather than an error. 3. That route now carries **all** the reads the tier had been absorbing, arriving together with no ramp. This is the **rebuild path**: the load the system of record takes while the tier is absent or refilling. 4. If the source cannot take that rate, the failure moves rather than disappearing. The tier's outage becomes the source's outage, and requests that never touched the tier start failing too. ## The facts that decide degrade or fail - **The role.** Derived copy: degradation is possible. Sole home: the fact is gone, and there is nothing to degrade to. - **The route.** Is the read to the system of record actually present on that code path, or does the code assume an answer comes back? A design with no route is not degrading; it is failing more slowly. - **Headroom at the source.** The unfiltered rate has to be survivable, or degradation is just a different failure. A source sized in the presence of the tier has never served this rate. - **How the tier stopped answering.** The four modes below produce four different caller experiences. - **How many services share the tier.** A shared tier's bad minute is shared: every service on it degrades in the same minute, and each one's system of record takes its own unfiltered rate simultaneously. ## Stopping answering is not one behaviour | How it stops | What the caller sees | | --- | --- | | Unreachable | Nothing returns; the request waits, then fails, on every call that consults the tier | | Answering slowly | Every consulting request pays the delay and holds its resources while it does | | Emptied by a restart | Every read is a miss; answers stay correct while the full read load lands on the source | | At its memory ceiling | Depends on the store: some reclaim entries to accept writes, others refuse writes and keep serving reads | That last row is where confident answers most often go wrong. There is no single ceiling behaviour in this class of store, so a design that survives only one of them is half written. The same caution applies to restart: what a store keeps across one varies, and the caller cannot assume a populated tier afterwards. ## Why slow can be worse than absent A component that refuses immediately at least releases the caller. A component that answers eventually keeps a thread and a connection occupied on every request while it does, so the service's own capacity drains from the inside even though nothing has reported an outage. This is why "is it up?" is the wrong monitoring question for a tier on the hot path, and why an answer that treats the outage as binary misses the mode that actually takes services down. ## What a strong answer states explicitly Name the role, name the route, name who takes the load, and say what the user sees at the end: a slower page, a page missing a section, or an error. "It just reads the system of record instead" is only true if that branch exists in the code and the source can take the traffic; otherwise it is an assumption doing the work of a design. The candidate who has operated one of these volunteers the rebuild-path load without being asked, because that is the part that turns a tolerable outage into an incident.
- How does the answer change if those entries are the sole home of ephemeral state rather than a derived copy?It stops being a latency question. There is no system of record to route to, so the facts the entries recorded are gone: whatever they represented has to be treated as lost, and the design has to say what the service does with the loss rather than how it serves the read more slowly.
- The tier is reachable but answering slowly. Why can that be worse for the caller than it being unreachable?Because every consulting request keeps a thread and a connection while it waits, so the service's own capacity drains even though the tier reports as up. A component that refuses immediately at least frees the caller to do something else.
- What should have been measured in advance to know whether the system of record survives the unfiltered rate?The read rate the tier is absorbing and the rate the source has actually served. If the source was sized after the tier was in place, those two numbers were never the same, and the gap between them is the size of the incident waiting to happen.
saying these in an interview costs you the question
- If the tier is down the service just reads the source; nothing changes.
- An unreachable component fails instantly, so requests are not slowed.
- The source can obviously take it; it used to serve all the traffic.
- A tier that is up is a tier that is working.
- Only the service that wrote the entries is affected when a shared tier fails.