skip to content

A tier introduced as a derived copy now holds facts nothing else can produce. What has changed, and how do you get out?

level: principalimportance: should knowfreq 44%

answer

  1. no code changed; the invariant did
  2. availability became durability
  3. routine pressure can delete a fact
  4. maintenance became migration
  5. exit per class, not wholesale

basics

~20 s

No code changed; what changed is where those facts live. The tier's availability is now the service's durability, routine memory pressure can remove a fact nothing rebuilds, and resizing it became a data migration. Exit per entry class, never wholesale.

solid answer

~50 s

The software may be identical to the day it shipped; what moved is the invariant. Four things follow. The tier's availability has become the service's durability, so its failure modes are now data-loss modes. On stores that reclaim entries under memory pressure, the store may remove the only copy of a fact as a routine response to a full memory ceiling, and it is behaving correctly when it does. Operations that were once free — resizing, replacing, moving, running at a smaller ceiling — are now data migrations. And the people who choose the memory ceiling are usually not the people who know the fact matters. The exit is per entry class: give the class a system of record and it reverts to a derived copy; or accept it deliberately as the sole home for ephemeral state, with the loss priced and written down; or remove the class. Making the store survive restarts is not the exit.

go deeper

for a junior

The takeaway is the habit, not the process: before you write something into a fast shared tier and nowhere else, ask who else will have it tomorrow. If the answer is nobody, raise it rather than deciding alone.

for a middle

Be able to explain why this is not a bug in the store. A tier configured to reclaim entries under memory pressure is behaving correctly when it removes one; what is wrong is that a fact nothing can rebuild was put somewhere with that behaviour.

for a senior

Show that you would triage per entry class and have an opinion on each: durable home, accepted loss, or delete. Be specific about what becomes risky in the meantime — resizes, replacements and moves are now data operations.

for a principal

The deliverable is the position, not the migration: state derivability as a reviewable invariant with an owner, price the exit per class, say plainly how long the tier has been load-bearing, and for a shared tier make the other teams' inherited risk visible to them.

## Nothing broke; an invariant moved This is the characteristic failure of a volatile tier, and it never arrives as an incident. It arrives as a series of reasonable small decisions: a counter that was easier to keep here than to write through; a flag someone set during a migration and never moved; a queue of pending work that only ever existed in memory; a feature shipped on a Friday by a team that had never been told the rule. Each one is defensible on its own. Together they change what the tier **is**, while every document still calls it a derived copy. The important observation for a design review is that the change happened in the **invariant**, not in the code. No commit says "this tier is now a system of record". That is why it is only ever discovered during an incident, or by someone who goes looking on purpose. ## What is now true, whether or not anyone has said so 1. **The tier's availability is the service's durability.** Every failure mode of a component chosen for speed is now a mode in which facts cease to exist. Nothing about the component was designed with that job in mind. 2. **Routine memory pressure can delete a fact.** On stores that reclaim entries when they reach their memory ceiling, the store may remove the only copy of something nobody can rebuild — and it is not malfunctioning, it is doing precisely what a volatile tier does. On stores that refuse writes at the ceiling instead, the same pressure shows up as writes failing, which loses the new facts rather than the old ones. Either way the loss is caused by ordinary operation. 3. **Maintenance became migration.** Resizing, moving, replacing or reconfiguring the tier used to be free because the contents did not matter. Each of those is now a change with a correctness risk and a rollback question. 4. **Authority is in the wrong hands.** The memory ceiling, the reclamation behaviour and the maintenance window are chosen by people operating a cache-shaped component. They have no way to know that one class of entry in it is a business record. 5. **In a shared tier, the blast radius is other teams'.** One team's drift changes what an outage means for every service on that tier, and none of the others will find out until it happens. ## The move that looks like an exit and is not The reflex is to make the store keep more across a restart and declare the problem solved. It is not an exit, for three reasons: what a store retains across a restart varies by store and configuration and comes with its own trade-offs to be chosen deliberately; it does nothing about entries removed under memory pressure while the process is happily running; and it leaves every operational move — resize, replace, relocate — still carrying data. The tier does not become a system of record because it survives a restart. It becomes one when somebody decides it should be, which is precisely the decision that never got made. ## The exit, taken one class at a time Migrating the tier wholesale is both unnecessary and unsafe, because the classes in it want different outcomes. For each class with nothing behind it, choose one: | Option | What it means | When it is right | |---|---|---| | Give it a system of record | Write the fact to a durable home; the tier reverts to a derived copy of it | The fact has value beyond the current request, and someone would notice it missing | | Accept it as the sole home for ephemeral state | Decide the loss is tolerable, write down exactly what breaks, and tell the operators | The fact is genuinely short-lived and its loss degrades rather than corrupts | | Remove the class | Delete the writer and the readers | The fact is a leftover, which is more common than anyone expects | The second option is a real answer, not a cop-out — much of what a volatile tier legitimately holds has no other home by design. What makes it acceptable is that it was chosen, priced and communicated, rather than discovered. ## Keeping it from happening again The repair that lasts is procedural, because the drift is procedural: - **State the invariant where it is enforced.** Every entry class in this tier names what regenerates it — a line in the design review checklist, not a paragraph in a document nobody opens. - **Give the tier an owner** who is asked that question when a new class appears, and who can say no. - **Re-ask it on a schedule**, because the answer decays as features are added by people who joined after the rule was written. - **For a shared tier, publish the classes**, so that other teams' understanding of an outage matches yours. ## The judgement this question is really testing A candidate at this level is expected to notice that the interesting question is not technical. The technical work — write the fact somewhere durable — is usually a few days. The hard parts are deciding which classes deserve it, admitting in writing that the tier has been load-bearing for months, and putting a rule in place that survives the people who wrote it. The answer that impresses names the drift, prices it, and describes the exit as a sequence of per-class decisions with an owner attached.

  • Why is accepting the loss sometimes the correct exit rather than an evasion?
    Because some state genuinely has no other home and short-lived degradation is an acceptable outcome. What separates a decision from an evasion is that it is written down: which class, what breaks when it goes, who is told, and what the operators may therefore do to the tier without asking.
  • How would you make the case for the work to people outside engineering?
    Not in terms of keys and memory. Name the fact that disappears, who notices, and what they cannot do afterwards — then say that the component holding it was chosen for speed and is operated on that basis, so the loss needs no failure, only a busy day.
  • What signal would tell you drift is happening again a year from now?
    New entry classes appearing with no answer to what regenerates them, and the tier's memory ceiling being discussed as a capacity number by people who do not know what is in it. Both are visible before an incident if someone is asked to look.
  • Does moving to a managed in-memory service change the picture?
    It changes who performs maintenance, not what the entries are. The operator is now a provider working to its own schedule and its own idea of what the component holds, which makes the unstated assumption harder to protect rather than easier.

A team keeps a printed copy of the register on the desk because walking to the filing room is slow. For months the printout is only a copy. Then somebody, in a hurry, writes a correction on the printout and not in the register. Nothing about the paper changed, and no one announced anything — but the desk copy is now the register, and the cleaner who throws it out on Friday will be doing exactly the job they were given. Laminating the printout so it survives being knocked off the desk does not make it the register either; it just makes losing it harder to notice.

saying these in an interview costs you the question

  • Thinks making the store survive restarts resolves the drift.
  • Treats the tier as one thing to migrate rather than a mixture of classes.
  • Calls the store faulty for reclaiming an entry it was configured to reclaim.
  • Presents accepting the loss as a cop-out rather than a priced decision.
  • Leaves the rule in a document instead of the review that enforces it.
  • Ignores that other teams on a shared tier inherit the new failure meaning.