skip to content

One in-memory tier holds both reconstructible copies and records that exist nowhere else: which ceiling posture do you choose, and what does the choice cost?

level: principalimportance: should knowfreq 44%

answer

  1. ask what a lost entry costs
  2. one setting, two kinds of contents
  3. loud and recoverable beats silent and lossy
  4. different loss semantics, different failure domains

basics

~20 s

Neither posture is right for both halves, so the real answer is to split them onto separate tiers. Forced into one, choose by blast radius: silent removal destroys the only-copy records, while refusal stops writes loudly but loses nothing already stored.

solid answer

~50 s

The deciding test is not the workload's size or its hit ratio — it is whether a removed entry can be reconstructed from a source of truth. Mixing both kinds under one memory ceiling means one posture must be wrong for one half: **removal** quietly destroys the records that exist nowhere else, and **refusal** denies writes even for the half whose loss would have been cheap. The principal answer is usually to stop mixing them, giving each class its own tier and therefore its own failure domain. Where one tier is forced, prefer the loud failure, because a refusal is recoverable and reported while a removal is neither. The reverse case is real, though: if losing some reconstructible state degrades gracefully and refusing writes halts the whole path, removal is the better trade. Note that not every store lets you choose a posture at all.

go deeper

for a junior

The idea to hold onto: losing an entry that exists nowhere else is data loss, while losing a copy is only a slower read. That difference is what the setting is really about.

for a middle

Explain that the posture applies to the whole tier while the consequence differs per entry, and that removal is silent while refusal is reported to the caller that hit it.

for a senior

Show the diagnosis and the in-tier mitigation: which entries are reconstructible, whether eligibility can protect the rest, and what the resulting failure looks like in production for each choice.

for a principal

Own the split. Argue for separate failure domains for copies and only-copy state, state where implementations give you no choice at all, and be able to argue the reverse case where graceful loss beats a stalled write path.

This is a design decision disguised as a configuration value. The setting is one line; the thing being decided is what your system does to itself when memory runs out, and for which half of its contents. ## The test that decides it One question separates the two kinds of entry in the tier: > **If this entry disappeared right now, could the system reconstruct it from somewhere else?** - **Yes** — it is a copy. Removing it costs a slower read while it is rebuilt. That is a performance event. - **No** — it is the only record of the fact. Removing it costs the fact. A session that no longer exists, a lease nobody holds, a counter that lost its history, a marker saying a job already ran. That is a correctness event, and no retry recovers it. Everything else — which entries are hot, how large they are, what the hit ratio is — is secondary to this, because the two answers have different *kinds* of consequence, not different magnitudes of the same one. ## The three postures by blast radius | Posture | Who finds out | What is lost | Recoverable? | |---|---|---|---| | Refuse writes | The caller, immediately | Nothing already stored | Yes — the write can be reissued | | Remove entries | Nobody, until a later read | The removed entries | Only if a source of truth exists | | Be killed | Everyone at once | The entire keyspace | Only from whatever durability exists | Read top to bottom, the trade is loudness against continuity. Refusal is the only one that keeps stored state intact and tells you it is unhappy. Removal keeps the system serving by paying in state. The kill is not a posture anyone chooses; it is what you get when the first two were never reachable. ## Why the mix is the actual problem With both kinds of entry behind one memory ceiling, the posture applies to the tier and not to the entry: 1. **Choose removal** and pressure will eventually take an only-copy record, silently. The system will not error; it will behave as though something that happened never happened. 2. **Choose refusal** and pressure stops writes for everything, including the half whose entries could have been dropped for the price of a rebuild. There is a partial in-tier answer where the store supports it: restrict removal to entries carrying a lifetime, and give a lifetime only to the reconstructible half. The permanent records are then protected by eligibility. The catch is the dead end at the other end of that road — when the eligible population empties out, the store has nothing to remove and refuses anyway, which is the refusal posture arriving without anyone choosing it. ## The answer a principal usually gives **Split the tier.** Copies and only-copy state have different loss semantics, so give them different failure domains: separate tiers, separate ceilings, separate postures. The copies tier can remove freely; the state tier refuses and is sized so that refusing is rare. The cost is honest and boring — a second thing to run, monitor and pay for — and it is usually smaller than the cost of the first silent loss. When one tier is genuinely forced, pick by which failure your system survives better: - **Prefer refusal** when the tier's state is load-bearing and its loss is not detectable by the application. A loud failure you can see beats a silent one you cannot. - **Prefer removal** when losing some entries degrades gracefully but a stalled write path does not — a tier where a missing entry means a slower path, and refusing writes means the whole request fails. Arguing this direction is legitimate; the mistake is arriving at it by default rather than by argument. ## The parts that are not yours to assume - **Not every store offers a choice.** Some in this class always remove and expose no refusal posture at all; on those, "choose the posture" really means choose the store, or choose the deployment. Defaults differ too, and the default is what most systems are actually running. - **Some stores have no ceiling concept**, in which case the operating system's out-of-memory kill is the only posture available, and the whole discussion becomes a question of headroom. - **What the application does with a refused write** — reissue, drop, route elsewhere — is a separate subject, and so is how the ceiling and its headroom are chosen. ## The interview signal A candidate who answers this well says three things: that the decisive question is whether a source of truth exists, that a posture applies to a whole tier while the consequence varies per entry, and that the honest resolution is usually to stop mixing the two. A weaker answer reaches for hit ratio, or names the posture they have used before as though it were the behaviour of stores in general.

  • Can you protect the only-copy half without running a second tier?
    Partly, where the store supports restricting removal to entries that carry a lifetime: give lifetimes only to the reconstructible half and the permanent records become ineligible. The catch is that once the eligible population empties, there is nothing left to remove and the store refuses writes anyway — the refusal posture arriving without anyone choosing it.
  • When is silent removal the right call even for state with no source of truth?
    When the loss degrades gracefully and the alternative does not. If a missing entry costs one user one retry while refusing writes fails every request on the path, removal is the smaller harm. The requirement is that you arrived there by argument, measured the loss, and can see it — not that it was the default.
  • The store we use offers no refusal posture at all. What then?
    Then the posture decision moves up a level: you are choosing a store, or a deployment, rather than a setting. Keep only-copy state off that tier, or accept that pressure means loss and design for it — detect the absence, rebuild what you can, and make the tier's capacity the control you actually manage.

saying these in an interview costs you the question

  • Decides the posture by hit ratio rather than by loss
  • Assumes every entry can be reloaded from a database
  • Treats removal as free because no error is returned
  • Believes one posture can be right for both halves
  • Says refusing writes is always the safe default
  • Presents their store's default as how stores behave