You inherit a shared tier that eight services read on their hot path; how would you decide whether it still earns its failure domain?
answer
- inventory by role, not by reader
- derived is removable; sole home is not
- eight services, one availability number
- spend the memory at the source instead
- derivability as a reviewable invariant
basics
~20 sInventory the entries by role, not by reader. Derived copies go if their sources can take the load; entries that are the only home of a fact cannot go at all. Then price eight services sharing one availability number.
solid answer
~50 sStart with what the tier holds. Entries that are **derived copies** are removable in principle: the question is whether each system of record can take its unfiltered rate. Entries that are the **sole home for ephemeral state** are not removable at all - removing them loses the facts, so that is an exit project with a design behind it, not a deletion. Then price what the shape itself costs: eight services on one component share one availability number and one memory limit, so one team's growth is felt by seven others. Compare against what you would spend instead - more memory or headroom at the sources, fewer reads per request, or a per-instance copy where a read tolerates divergence. The common outcome is not removal but shrinkage: keep the part that earns the failure domain.
go deeper
The idea to take away is that removing a component is not the reverse of adding it. What the entries are - a copy of something, or the only record of it - decides whether removal is even possible.
Be able to sort a keyspace by role and say, for each portion, where a read would be served from if the tier were gone. That inventory is the whole basis of the decision.
Demonstrate that you would gather evidence rather than argue: attribute the reads, compare each source's served rate against the unfiltered rate, and shed a controlled share at a quiet hour before committing.
Own the position: derivability stated as a reviewable invariant, the drift into an unacknowledged system of record priced with its exit, and the coupling of eight services weighed against the saving of operating one component.
## The question is about entries, not services Eight readers is a fact about blast radius, not about removability. What decides whether the tier can go is what its entries *are*. Sort the keyspace by role before anything else: - **Derived copy** - a faster copy of data a **system of record** still holds. Removable in principle; the cost of removing it is load and latency at the source. - **Sole home for ephemeral state** - facts with no other holder. Not removable. Getting out of this is a migration with a design behind it, and the honest cost of the tier includes that exit. - **Coordination point** - several processes agreeing, through one shared place, on who holds something. Removing it is a change to how those processes agree, which is a correctness change, not an infrastructure change. - **Transient transport** - delivery to whoever is listening now. Removing it changes who hears what, and anything that needs retention belongs elsewhere anyway. The inventory usually surprises people: a tier introduced years ago as one role has typically accumulated all four, each added by a different team on a different day. ## Evidence you can gather without turning it off 1. **Attribute the read traffic by prefix and by service**, so each portion can be argued about separately rather than as one aggregate. 2. **Establish, per portion, whether a rebuild path exists at all** - can this entry be recomputed or re-read from somewhere authoritative? If nobody can say, that is your finding. 3. **Compare the rate each source would receive against the rate it has actually served**, remembering that sources sized after the tier existed have never served the unfiltered rate. 4. **Shed a small, controlled share of reads past the tier at a quiet hour** and watch the source. This is the only evidence that is not an estimate. 5. **Watch what an ordinary incident already tells you.** A restart, a failover or a deploy that emptied the tier is a free experiment that has probably already happened; find it in the history. ## What the shape itself costs One tier read by eight services concentrates two things that were previously separate: - **One availability number.** Its bad minute is everyone's bad minute, and each of the eight sources takes its unfiltered rate in that same minute. - **One memory limit.** One team's growth in entries is felt by the other seven, and stores differ in how that ends: some reclaim entries to accept the write, others refuse writes while continuing to serve reads. Both are somebody else's outage. Against that, sharing genuinely saves: one component to operate, one upgrade, one set of alerts. The judgement is whether that saving is worth coupling eight failure stories together, and the answer changes with how critical the least careful of the eight services is. ## What you would spend instead | Alternative | What it buys | What it costs | | --- | --- | --- | | More memory or headroom at the sources | The source answers warm, so the reads the tier absorbed get cheap at their origin | Money, and a ceiling of its own | | Fewer reads per request | Removes the work rather than relocating it | Application change, sometimes a large one | | A per-instance copy inside each service | No hop, no shared component | Copies can diverge, and there is nothing to reconcile them - its own comparison | | Keep the tier, hold less in it | Keeps what earns the dependency | Requires the inventory above, and someone to keep it true | ## The invariant a principal states The reviewable rule is **derivability**: for every class of entries in the tier, someone can name where it would be rebuilt from. Not a habit, a stated invariant, checked when a new use is added. The tell that it has lapsed is procedural rather than technical: the runbook for an empty tier says *restore it* instead of *it refills*. At that moment the tier has quietly become a system of record - one with no durability posture anyone chose, no backup anyone tests, and an exit price nobody has costed. Pricing that exit is part of pricing the tier. ## The usual answer Rarely "remove it" and rarely "keep it all". It is normally: keep the portion whose reads repeat and whose source reads are expensive; move the sole-home portion to something that is meant to hold facts; and stop the tier from being the default destination for any state that has nowhere else to live. That is a smaller failure domain and a defensible one - which is what the question was actually asking.
- What is the tell that a tier has drifted from a derived copy into an unacknowledged system of record?Nobody can name the rebuild path for some class of entries, and the runbook for an empty tier says restore it rather than it refills. At that point the tier holds facts with no durability posture anyone chose, and the exit is a migration rather than a configuration change.
- Is the answer ever to keep the tier but shrink what it holds?Usually. Keep the entries whose reads repeat and whose source reads are expensive, move anything that is the only home of a fact to something meant to hold facts, and stop new state landing there by default. The failure domain stays, but it becomes one you can defend.
- How would you price the sharing itself, separately from the entries?By asking what the least careful of the eight services can do to the other seven: exhaust the memory limit, drive the read rate, or make a change nobody else reviewed. Sharing buys one thing to operate and sells the independence of eight failure stories.
saying these in an interview costs you the question
- Every entry is a derived copy, so the tier can simply be removed.
- A shared tier is cheaper because eight teams split one component.
- It has never gone down, so the risk is not worth pricing.
- The exit cost is a deployment change, not a design project.
- Removing the tier is always the honest answer once traffic drops.