What headroom and lead time should a growth policy for a shared in-memory tier carry, and why not wait until it is full?
answer
- full is already too late
- the move consumes capacity too
- trigger derived from lead time
- ceiling posture prices the overshoot
- shared headroom needs an owner
basics
~20 sEnough spare room to execute the move and enough lead time to finish it. Both scaling actions consume capacity while they run, so a tier that has already reached its ceiling is removing entries or refusing writes before any move can land.
solid answer
~50 sWrite the policy against three numbers: the growth rate you actually observe, the worst-case lead time of the move you would make, and what that move itself consumes while running. The trigger is the memory ceiling minus growth over that lead time minus the move's transient — usually much lower than teams expect, because acquiring capacity, moving key ownership, or sequencing a version change are measured in hours and days, not minutes. The policy also has to state which behaviour this store has at the ceiling, because that decides what overshoot costs: removal under pressure degrades quality quietly and buys you time, while refusing writes turns a missed trigger into caller-visible failures at once, and some stores in this class refuse by default. On a shared tier, add a rule about who may grow into the headroom, or the first owner to expand will spend everyone's margin.
go deeper
Understand that capacity for this tier is planned ahead of the shortage, because adding a node or a bigger machine takes real time and the tier behaves badly once it reaches its ceiling.
Be able to explain what happens at the ceiling — entries removed under pressure, or writes refused — and why that makes the distance to the ceiling a scheduling question rather than an alarm threshold.
Derive the trigger from the slowest move you actually have available, including caller-side changes, and show that you planned against the ceiling the server enforces rather than the host limit.
Own the policy across owners and time: who may grow into the shared margin, what the organisation accepts when the trigger is missed, and whether a tier with no move short of a caller release is a design you are prepared to keep.
Capacity policy for a volatile tier is not the same exercise as for a durable store, for one reason: the move that adds capacity needs capacity to run, and the state you are trying to avoid is a state in which the tier is already damaging data. "Scale when it is full" is not a conservative policy, it is a policy that guarantees an incident. ## Why "full" is too late At the memory ceiling the server does one of a small number of things, and which one it does is a property of the store and its configuration rather than a universal: - **removal under pressure** — it makes room by removing entries, ordered by recency, by frequency, arbitrarily, or only among entries that carry a deadline. Callers see a quality change rather than an error, which is a gentler failure but also a silent one; - **refusing the write** — it rejects writes and keeps what it holds. This is caller-visible immediately, and some stores in this class take this posture by default; - **neither, and the process is killed** — where the ceiling the server enforces sits above what the host or container will actually allow. Every one of those arrives before your new node does. Worse, two of the three make the move harder while it runs: a node under pressure removal is spending time and processor cycles removing entries, and a node refusing writes is generating caller errors you will be reading during the change. ## The three numbers a policy needs 1. **Observed growth rate** — how much stored-data size gains per week under current traffic, taken from the trend rather than the peak day. 2. **Worst-case lead time** — the full elapsed time of the move you would make, end to end. 3. **The move's own consumption** — the transient a change costs on the nodes that stay. Then the trigger is: `ceiling − (growth rate × lead time) − move transient`. Re-derive it when the growth rate changes; a threshold written once and never revisited is why tiers get caught. ## Lead time is the number people underestimate | move | what the elapsed time actually includes | |---|---| | bigger node | approval and procurement, standing up the replacement, moving callers, the interruption | | adding nodes / moving assignments | planning the target layout, the transfer itself under a throttle, verification | | version change | preconditions, one node at a time with a checkpoint between, the write-role changeover | | application-side change | releasing callers that address the tier differently, on their own release cadence | The last row is the one that turns days into weeks. If relieving the shortage requires callers to change how they address the tier, your lead time is somebody else's release schedule. ## Which numbers the policy is written against Write the trigger against the **ceiling the server enforces on itself**, because that is the bound whose breach changes behaviour. Keep the **host or container limit** as a separate alarm, because the gap between the two is what protects you from the process being killed outright, and the **resident footprint** as a third, because a footprint well above stored-data size means the process is holding memory it is not storing data in — which shrinks the real distance to the host limit without changing the ceiling at all. ## On a shared tier, headroom is contested A tier with several independent owners has one ceiling and one blast radius, so the spare capacity is a shared asset even though nobody bought it deliberately. Two consequences follow. First, unallocated headroom is consumed by whichever owner grows first, and the cost of that lands on everyone — including the owner whose writes are refused next, who may not be the owner who grew. Second, the policy has to name someone who can say no: either each owner carries a budget and is held to it, or the headroom has an owner who approves growth into it. "We will ask people to be careful" is not a policy, and neither is a prefix convention, which accounts for usage without bounding anyone's damage. ## What a strong answer includes - the store's ceiling posture, stated rather than assumed, because it prices the overshoot; - a trigger derived from lead time rather than from a round number like ninety percent; - the observation that a bigger node and a move of assignments have different lead times, so a tier that can only do one of them has only that one's trigger; - a rule for who may consume the margin on a shared tier; - and an honest statement of what the organisation does if the trigger is missed anyway — because the answer differs entirely between a store that quietly removes entries and one that starts failing writes.
- How do you turn "how long a move takes" into a memory threshold?Take the growth rate you actually observe, multiply it by the worst-case end-to-end lead time of the move you would make, add what the move itself consumes while running, and subtract the total from the ceiling the server enforces. The remainder is the trigger. Re-derive it whenever the growth rate or the available move changes.
- Does the same policy hold if the store refuses writes at the ceiling instead of removing entries?The trigger is the same; the cost of missing it is not. Removal under pressure degrades quality quietly and leaves you time to act, while refusing writes produces caller-visible failures immediately — and on a shared tier it fails whichever owner writes next rather than the one who consumed the room. Know the posture before setting the margin.
- Who should own the headroom on a tier shared by several applications?Someone who is able to refuse. Spare capacity that is not allocated gets consumed by whoever grows first, and the consequences land on every occupant. Either give each owner an accounted budget and hold them to it, or make the margin an operational asset with a named approver; a prefix convention alone accounts for usage without bounding damage.
saying these in an interview costs you the question
- Scale when the tier reaches its ceiling; earlier is wasted spend
- Headroom only has to cover growth, not the move itself
- Overshoot is free because the tier just removes the least useful entries
- The host memory limit is the number to plan capacity against
- On a shared tier, spare capacity belongs to whoever reaches it first