A team moves its tier to a different in-memory store; which part of its capacity model must be re-derived, and why?
answer
- split data facts from store facts
- counts travel, the constant does not
- everything downstream is re-derived too
- provenance, sensitivity, re-derivation trigger
- a stale constant fails silently
basics
~20 sThe fixed per-entry cost, and everything computed from it. Entry counts, payload distributions and growth rates describe your data and transfer intact; the per-entry constant describes the store's entry representation, lookup structure, allocator and build, and transfers not at all.
solid answer
~50 sSplit the model into facts about your data and facts about the store. Entry counts, the payload size distribution, the mix of entry shapes and the growth rate are yours and carry over unchanged. The **per-entry constant** is not yours: it is a property of the new store's header, its lookup structure's spare capacity, its allocator's rounding step, and whether a deadline costs anything there. So is everything derived from it — projected data size, the entry count at which the ceiling is reached, the headroom argument. Two second-order things also need re-deriving: whether the new store exposes an accounting of what it attributes to entries at all, and whether it has a ceiling concept of its own or leaves exhaustion to the operating system. The principal move is to make the constant an explicitly dated, re-measurable input rather than a number inlined in a spreadsheet.
go deeper
The idea worth carrying away is that the cost of an entry belongs to the store, not to your data. Two stores holding exactly the same keys and values can need noticeably different amounts of memory.
Be able to say which inputs are yours and which are the store's, and to name what actually makes the constant store-specific — header contents, the lookup structure's spare capacity, the allocator's rounding step, the cost of a deadline.
Demonstrate that you would re-measure rather than re-use, and that you would follow the change downstream into every projection built on the constant. Mention the second-order changes: whether per-entry accounting exists at all, and whether the store has its own ceiling.
Argue about what kind of object a capacity number should be. Provenance, a sensitivity band, a named re-derivation trigger and a production signal that contradicts the model early are the deliverables; the number itself is the least durable part of the work.
## Separate what the model knows about you from what it knows about the store A capacity model for an in-memory tier is, at bottom, one multiplication: entries times cost per entry, read against a ceiling. That makes it easy to audit, because every input is clearly owned by one side or the other. | Input | Owned by | Survives a migration? | |---|---|---| | Projected entry count | Your workload | Yes | | Payload size distribution | Your workload | Yes | | Mix of entry shapes | Your workload | Yes | | Growth rate | Your workload | Yes | | Fixed cost per entry | The store | **No** | | Projected data size | Derived | No, it is downstream | | Entry count at the ceiling | Derived | No, it is downstream | The asymmetry is the answer. Nothing about your data changed when you changed stores. Everything about what holding it costs did. ## Why the constant is a property of the store The fixed part of an entry's cost is assembled from implementation choices: - the contents and size of the entry header; - how much spare capacity the lookup structure carries so that it stays less than full; - the step the allocator rounds each of the several allocations per entry up to; - whether a deadline lives in a field that exists anyway or in a separate structure; - whether small values are placed inside the descriptor or reached through a reference. Every one of those differs across stores in this class, and several differ across versions and builds of the same store. That is why the constant does not travel — and why a constant carried across an upgrade is a smaller version of the same mistake. ## The second-order things that also change A migration can invalidate the model's *shape*, not just its numbers: 1. **Whether per-entry accounting exists.** Some stores report in detail what they attribute to entries; others report little more than a process total. Where the detailed figure is missing, the delta measurement still works, but on a coarser number — and the model has to say so, because the constant it produces now quietly includes things that are not entries. 2. **Whether the store has a ceiling of its own.** Where it does, the projection is read against a configured number the store acts on. Where it does not, exhaustion is decided outside the process, and the model's failure mode changes from a behaviour you configured to a process that disappears. 3. **Whether a lifetime is free.** If the old store charged nothing for a deadline and the new one keeps a separate structure for entries that have one, a workload where most entries carry a lifetime gets a step change that no payload arithmetic predicts. ## Making the constant auditable rather than remembered The practical principal move is not producing a better number. It is changing what kind of object the number is: - **Label it with provenance** — store, version, entry shape measured, sample sizes, date, and a description of how to re-run the measurement. - **State a sensitivity, not a point** — "at the measured constant we reach the ceiling at 21 million entries; if the constant is half again as large, at 14 million." A model that only produces one number cannot tell a reviewer how much the constant matters. - **Name a re-derivation trigger** — a version upgrade, a change in entry shape, a migration, or a calendar interval, whichever comes first. - **Close the loop with production** — compare the modelled cost per entry against what production currently demonstrates, and treat a drift as evidence the model has aged rather than as noise. ## The blast radius of a stale constant This matters because a stale constant is silent. Nothing fails while the tier is at a third of its ceiling; the model is simply wrong in a way nobody can see. The error surfaces at the ceiling, and by then the available moves are all expensive: enlarge the tier under load, change what the store does at the ceiling, or shed entries. If the tier holds state that exists nowhere else — claims, leases, records of work already done — shedding entries is not a miss to be re-read from somewhere, it is loss. A capacity model whose weakest input is undated and unowned is what converts a foreseeable arithmetic problem into that situation. ## The argument on the other side It is legitimate to carry a constant forward provisionally, and a good answer says when: during an evaluation, where the point is a rough comparison rather than a commitment; or where the tier runs so far below its ceiling that even a large error in the constant changes nothing. What is not legitimate is carrying it forward **silently**. The distinction a principal is expected to draw is between a rough number that is labelled rough and a rough number that has been quietly promoted to a fact.
- Which inputs to the model genuinely do carry across the migration?The ones that describe your workload rather than the store: projected entry count, the payload size distribution, the mix of entry shapes, the growth rate, and the business requirement the tier exists to serve. None of those changed because the store did. Keeping them and re-deriving only the store-owned inputs is also what makes the migration's capacity work tractable rather than a rewrite.
- The new store exposes no detailed accounting of what it attributes to entries. Can the model still be built?Yes, with a stated caveat. The delta method works against whatever total the store does report: take a reading, write a large realistic sample, read again, divide. The resulting constant includes anything else that grew during the write, so it is an upper bound rather than a clean per-entry figure. Say that explicitly in the model, because an upper bound used as a point estimate is conservative in one direction and misleading in the other.
- How would you know the constant has aged without waiting for the ceiling to tell you?Compare it continuously with what production demonstrates: reported data size divided by reported entry count is a measured cost per entry that the running system produces for free. Alert on it drifting from the modelled value by more than the model's stated sensitivity. That converts an assumption into something the system contradicts early, rather than something the ceiling contradicts late.
saying these in an interview costs you the question
- Carries the old store's per-entry constant into the new model unchanged.
- Re-derives the constant but leaves the projections computed from it untouched.
- Assumes every store in this class reports what it attributes to entries.
- Assumes every store in this class has a configured memory ceiling of its own.
- Treats removal under pressure as harmless because entries can be re-read from elsewhere.
- Records a single projected number with no sensitivity and no re-measurement trigger.