How do you arrive at the number you set as an in-memory store's memory ceiling on a 32 GB node?
answer
- measured, not remembered
- entries touched in a named window
- ceiling below the limit, gap itemised
- alert before the ceiling, not at it
- name what would make you revise it
basics
~20 sMeasure what one real entry costs, multiply by the distinct entries touched in a named window, and set the ceiling there — far enough below the machine limit that the named headroom consumers fit in the gap.
solid answer
~50 sFour steps, in order. **Measure** the cost of one entry by writing a sample of real entries and dividing the increase in what the store attributes to its entries by the count — never quote a figure from memory, because it depends on the data and on the store. **Count** the distinct entries touched in a window you name, which is the working set; where the tier holds the only copy of something, count all live state instead. **Multiply** to get a subtotal and add the growth you expect before the next resize. **Subtract headroom** from the machine limit for replication buffers, client output buffers and duplication while a whole-keyspace copy runs, and set the ceiling below what is left. Then alert below the ceiling with lead time to act, and re-derive the number when the miss rate climbs at flat traffic or the resident size drifts from the store's own accounting.
go deeper
Know that the number is arrived at by arithmetic: how many entries will be touched, multiplied by what one entry costs, kept under the memory the machine has. Both factors are measured, not remembered.
Walk the steps in order and say where each figure comes from. The part most often skipped is measuring cost per entry on a real sample instead of quoting one, and naming the window the working set was counted over.
Show the instrumentation, not just the arithmetic: an alert below the ceiling with lead time, a second one on the process's resident size, and a written list of the conditions that would make you re-derive the number.
Treat the ceiling as a decision with an owner and a review trigger, and say out loud where implementations diverge — whether a ceiling exists, what it counts, what an entry costs — rather than presenting one store's arrangement as the method.
## The procedure The number you set is an argument, not a recollection. It is defensible if you can produce it in this order: 1. **Measure the cost of one entry on your own data.** Load a sample of real entries into a store of the same kind and configuration, take the increase in what the store attributes to its entries, and divide by the number written. A thousand representative entries is enough to be useful, and the mix has to be representative — a sample of your smallest values yields a figure that flatters everything above it. 2. **Count the working set.** Distinct entries touched in a window you name out loud, chosen to include the busiest period you are sizing for. Where the tier holds the only copy of something — sessions, leases, quotas — that content is not sized by what is touched but by what is live, because an entry you do not hold is gone rather than re-read. 3. **Multiply, then add planned growth.** Entries multiplied by measured cost per entry, plus whatever you expect to be added before the next time anyone revisits this number. Growth belongs in the ceiling, not in the headroom; they are different arguments. 4. **Fit it under the machine limit with the headroom named.** The gap between the ceiling and the machine limit holds memory the process carries on behalf of things that are not entries: buffers for followers, buffers for connections that have not drained their replies, and transient duplication while a whole-keyspace background copy is written. Each item is included only if the deployment actually has it, and each is sized by what drives it rather than as a share of the total. If the subtotal plus the headroom does not fit under the machine limit, that is the answer to a different question — whether to hold less, or to run on more — and it should be raised as such rather than absorbed by shaving the ceiling. ## Measure, do not estimate The per-entry cost is the factor with the widest spread and the one most often guessed. The same logical value costs different amounts on different stores and at different value sizes, because the bookkeeping kept per entry, the layout of the value and the way the allocator rounds a request all differ. A remembered figure carries somebody else's data shape and somebody else's store. The measurement also happens to answer questions the arithmetic cannot: it exposes the entries that are far more expensive than the average, and it tells you what the store's own accounting reports, which is the number your ceiling will be compared against. ## The number is a hypothesis A ceiling set once and never revisited is a ceiling that will eventually be wrong, because both factors move. Name the triggers in advance: - **The miss rate climbs at flat traffic** — the working set has outgrown what the ceiling holds, on one of its two axes. - **The store's accounting per entry drifts upward** — the payload changed, and the entry count is no longer a proxy for memory. - **The process's resident size pulls away from the store's accounting** — memory is being held outside the entries, which is a statement about headroom rather than about the ceiling. - **A consumer of the headroom is added** — a follower, a second region's feed, a newly enabled whole-keyspace copy. The gap has to grow before the feature is switched on, not after. And instrument it so those triggers arrive early: the alert belongs **below** the ceiling, with enough lead time to add memory or reduce what is held, not at the ceiling, where the only remaining options are the ones the store takes by itself. ## Where implementations diverge This is the part that separates a sizing argument from a recipe: | Dimension | How stores differ | What it does to the procedure | |---|---|---| | Is there a ceiling at all? | Some expose a configurable one; on others the machine or container limit is the only bound | With no setting, step 4 moves into the node choice and the margin is watched rather than enforced | | What the ceiling counts | Implementations draw the line around their own bookkeeping and buffers differently | Verify the gap by measuring the process's resident size under real traffic, rather than trusting what you believe the setting excludes | | Cost per entry | Varies by store, by value shape and by value size | Makes step 1 non-optional and makes any quoted figure unusable | | Headroom consumers present | No followers means no replication buffers; some stores never write a whole-keyspace copy | Headroom is itemised per deployment, never a blanket percentage | ## The shape of a defensible answer An interviewer is listening for four things: that you measured rather than recalled, that you said which window your working set was counted over, that the ceiling sits below the machine limit with the gap justified item by item, and that you named what would make you change the number. A candidate who produces a single figure with no window, no measurement and no gap has described a setting, not a decision. ## One worked arrangement, not a formula: the ceiling is the working-set subtotal, the three headroom lines sit between it and the machine limit, and the remaining margin absorbs the difference between the store's own accounting and what the process holds. Every line is a measurement on this deployment — the headroom lines are included only where the deployment has that consumer, and the cost per entry is measured rather than quoted ``` measured cost per entry (sample of 1,000 real entries) = 320 bytes distinct entries touched in the busiest hour = 72,000,000 --------------------------------------------------------------------------- working-set subtotal 320 bytes x 72,000,000 = 23.04 GB headroom, replication buffer for one follower = 1.50 GB headroom, client output buffers at peak connection count = 1.00 GB headroom, duplication while a whole-keyspace copy runs = 3.00 GB --------------------------------------------------------------------------- ceiling to set (the subtotal) = 23.04 GB ceiling + headroom = 28.54 GB machine limit = 32.00 GB remaining margin = 3.46 GB ```
- The store you are sizing exposes no ceiling setting at all. What changes?The enforcement point moves to the operating system, so the margin has to be larger and watched rather than configured. The same arithmetic chooses the node instead of the setting: working set plus itemised headroom decides the machine you run on, and the guard rails become an alert on the process's resident size and a bound on what the workload is allowed to write.
- Where should the alert threshold sit, and against which number?Below the ceiling, on the store's own accounting, far enough below that someone can act before the store has to. Add a second alert on the process's resident size against the machine limit, because buffers and a running copy move that number without moving the first one at all.
- A colleague proposes reserving 50% of the node as headroom. What is wrong with that?It is not wrong so much as unaccountable. A percentage cannot be checked, cannot be reduced when a consumer is absent, and cannot be grown when a real one is added. Itemised headroom answers all three: no followers means that line is zero, and enabling a whole-keyspace copy means a specific line grows before the feature is switched on.
saying these in an interview costs you the question
- Quotes a per-entry overhead figure from memory instead of measuring
- Gives a ceiling with no window named for the working set
- Reserves a round percentage of the node with no named consumer
- Alerts at the ceiling rather than well below it
- Treats the number as settled once configured
- Assumes every store offers a ceiling setting to configure