In one self-run in-memory store holding both derived copies of catalogue rows and the only record of in-flight claims, how should restart survival, entry lifetime and sizing differ between the two populations?
answer
- the role, not the component, sets these
- loss window versus slow period
- deadline: freshness bound, or promise
- size by hit rate, or by arrival times lifetime
- strictest population sets the posture
basics
~20 sDerived copies need no restart survival, take a freshness-bound deadline, and are sized to the hot subset. The claims need restart survival or a stated loss window, deadlines that carry meaning, and sizing for the whole live population.
solid answer
~50 sThe role decides all three. **Derived copies**: an empty start is a slow period, not a loss, so restart survival is optional; the deadline is a freshness bound and may be shortened for capacity; sizing targets the hot subset, and undersizing only means more reads reach the system of record. **The claims**: nothing else holds them, so either the store keeps them across a restart or the design accepts a stated loss window; their deadline is part of what the system promises and cannot be trimmed for capacity; sizing must cover the whole live population with headroom, because an entry that is not held is a lost fact rather than a slow read. That also means the two populations disagree about what should happen at the memory ceiling, and stores in this class differ in whether they reclaim entries there or refuse further writes.
go deeper
Notice that two kinds of entries can live in the same store and need opposite treatment. A copy of something the system of record still holds can be thrown away; a record nothing else holds cannot.
Be able to derive the three settings from the role: what an empty start means, what the deadline is a statement about, and whether sizing comes from a hot subset or from arrival rate multiplied by lifetime.
Demonstrate that you confirm what this store does at the ceiling and across a restart rather than assuming it, and that you can state the loss window you accept for the population that has no other holder.
Own the decision that one deployment runs at the strictest population's posture. Price what that costs the cheap population, and say at what size you would separate them instead of carrying the cost quietly.
## Two populations, one component The deployment is one **in-memory store** run by the team itself. The entries in it fall into two populations with two different roles: catalogue rows are a **derived copy** of data a **system of record** still holds, and in-flight claim records are the **sole home for ephemeral state** - nothing else in the system has them. Every setting worth arguing about follows from that distinction rather than from the component. ## The derived copies - **Restart survival**: optional. An empty start means callers read the system of record instead for a while. That period has a cost, but it is a latency and load cost, not a loss of information. - **Entry lifetime**: the deadline is a **freshness bound** - a statement of how stale this copy may be allowed to get. It is a design knob, and shortening it to fit more entries changes only freshness, not correctness. - **Sizing**: the hot subset. Undersizing shows up as more requests reaching the system of record, which is visible and gradual. - **At the memory ceiling**: reclaiming entries under memory pressure is acceptable, because a reclaimed copy can be read again from its authoritative holder. ## The only record of in-flight claims - **Restart survival**: this is now a real decision. Either the store keeps this population across a restart, or the design states the **loss window** it accepts and what the application does with claims that vanish. Not every store in this class offers restart survival at all, and on those that do it is a posture with a running cost - which is why the decision belongs in the design rather than in a runbook. - **Entry lifetime**: the deadline is part of the behaviour. It bounds how long a claim stays in flight before the system treats it as abandoned, so it is chosen from the business process it represents and cannot be trimmed to save memory. - **Sizing**: the **whole live population plus headroom**, computed from the arrival rate and how long an entry lives, not from a hit ratio. There is no such thing as a miss here; there is only an entry that exists or a fact that is gone. - **At the memory ceiling**: reclaiming is silent loss. For this population, a store that refuses further writes at the ceiling is failing loudly and correctly, which is the better failure - provided the application is written to handle a refused write. ## Side by side | Setting | Derived copies of catalogue rows | The only record of in-flight claims | |---|---|---| | Empty after a restart | Slow period; reads go to the system of record | Claims are gone; someone must be told | | Meaning of the deadline | How stale a copy may get | How long a claim stays in flight | | May the deadline be shortened for capacity? | Yes, it changes freshness only | No, it changes what the system promises | | Basis for sizing | Hot subset and hit expectations | Arrival rate multiplied by lifetime, plus headroom | | Reclaimed under pressure | Acceptable | Silent data loss | ## Where stores differ, and why that matters here It is tempting to state ceiling and restart behaviour as properties of in-memory stores. They are not properties of the class: - Some stores **reclaim** entries when memory runs short; others **refuse** further writes; a badly bounded one is simply killed by the operating system. Which of these your store does is something to confirm, not assume. - Some stores keep **nothing** across a restart by design, some offer restart survival as a choice with a cost, and the size of the loss window varies with the posture chosen. - Some acknowledge a write only once another node also holds it, while others acknowledge as soon as the node taking the write has it. That difference is invisible for the catalogue copies and material for the claims. What is universal is the reasoning, not the behaviour: name the role, then confirm what this particular store does at the ceiling and across a restart. ## If they must share one deployment They can, and the arithmetic is simple and unpleasant: **the strictest population sets the posture for the whole deployment**. If the claims must not be reclaimed, the deployment cannot be configured to reclaim freely, so the catalogue copies lose the property that made them cheap - you must now size for the union rather than for the hot subset, and you carry restart survival for entries that never needed it. That is a real cost, and it is the argument for separating the two populations onto different deployments once either is large enough to notice. The decision is not "one tier or two" as a matter of taste; it is whether the derived copies are cheap enough to run under the claims' rules.
- How do you size the claim population when there is no hit ratio to reason from?From arrival rate multiplied by how long an entry lives, plus headroom for the worst burst you expect and for growth between capacity reviews. There is no miss to absorb the error: an entry that does not fit is a fact that is gone, so the margin is a correctness margin rather than a performance one.
- Why is a store that refuses writes at its memory ceiling sometimes the better behaviour?Because for a population with no other holder, reclaiming is silent loss while refusal is a visible error the caller can handle, retry or escalate. It is only better if the application actually handles a refused write; if it ignores the error, the loud failure becomes a quiet one again.
- What does the claims population cost the catalogue copies when they share one deployment?The strict settings apply to everything: no free reclamation, restart survival carried for entries that never needed it, and sizing for the union rather than the hot subset. The copies lose the cheapness that justified them, which is the usual reason to split the populations once either grows.
saying these in an interview costs you the question
- Sizes a population with no other holder from a hit ratio.
- Shortens a deadline to save memory without asking what it means.
- Assumes the store reclaims entries when memory runs short.
- Calls an empty start after a restart a loss for every population.
- Runs both populations under one posture without pricing it.