You own a shared in-memory store used by many teams — what rule would you set for what one entry may hold?
answer
- convention, not configuration
- two numbers: bytes and members
- the store's ceiling is a backstop
- some data does not belong here at all
basics
~20 sA stated bound in bytes and in members, derived from the latency and memory you are willing to spend on the worst operation rather than from whatever ceiling the store enforces, applied on the write path and re-examined whenever a new growth shape appears.
solid answer
~50 sThe store has no schema, so nothing but convention decides what one key holds, which makes this an ownership question rather than a configuration one. State the rule as two numbers, a byte bound and a member bound, because they fail differently: bytes bound what an operation moves, members bound what a whole-entry operation costs in time. Derive both from what you are willing to pay when something touches that entry, not from the store's own per-entry size ceiling — where a store enforces one it sits far above the size that already hurts, and the numbers differ by orders of magnitude between stores. Enforce the rule where writes happen, in a shared wrapper every team uses, and keep an exit clause: data that must keep everything forever does not belong in an unschematised keyspace at all.
go deeper
Recall that nothing in this kind of store limits what one entry holds by itself, so any limit that exists is one a team agreed and wrote into its own code.
Explain why a bound needs a member figure as well as a byte figure, and why the size a store refuses is not a useful design input: it sits far above the size at which an entry is already slow to move and slow to free.
Show where enforcement lives on a tier that is already running: a shared write path that refuses or trims, a measurement that reports entries nearing the bound, and a plan for the shapes that are already over it when you introduce the rule.
The call is organisational. Decide who owns the keyspace of a shared tier, what the convention must encode for a breach to have a name attached, and when the honest answer is that this data belongs in a durable store instead of being bounded here.
## Why this is a rule and not a setting An in-memory store of this class has no schema. There is no declared type, no maximum length attached to anything, and nothing that objects while an entry grows. Where the store enforces a **per-entry size ceiling** at all, there is nothing between an ordinary entry and that ceiling. So what one entry may hold is not a property you configure; it is a convention somebody writes down, and the whole question is who owns it and how it is kept true. On a shared tier the stakes differ from a single service's. Teams do not see each other's keys. One team's entry that grew to hundreds of megabytes is paid for by everyone else: in the memory ceiling they share, in the worker time an operation on it occupies, and in the reclamation when somebody finally removes it. That is why the rule is worth having before anyone needs it. ## Two numbers, and where they come from 1. **A byte bound.** This bounds what one operation moves and what one key costs against the shared ceiling. It is the number that keeps a single entry from being a meaningful fraction of the tier. 2. **A member bound.** This bounds what a whole-entry operation costs in time. A collection of several million tiny members can look unremarkable in bytes and still be expensive to build a reply from, to rewrite, or to free. Bounding one does not bound the other, which is why a rule with only a byte figure keeps being breached by shapes nobody expected. Derive both from what you will accept when something touches that entry — a worst-case latency you are willing to publish, a share of the ceiling you will let one key hold — rather than from what the store permits. ## The store's ceiling is a backstop, not a bound | | the bound you choose | the ceiling the store enforces | |---|---|---| | derived from | your latency and memory budget | the store's own internal design | | where it sits | near the size that starts to hurt | far above it | | what happens at it | the write path refuses or trims | behaviour varies: some stores refuse the write, others accept whatever fits in memory | | travels to another store | yes | no | Quoting a particular ceiling figure is product knowledge rather than design, and it is the fastest way to give an answer that is wrong for whoever is interviewing you. What travels is the reasoning: a ceiling exists for the store's own reasons, it is not a statement about what your workload should do, and the number you design against is always the smaller one you chose. ## Where a rule like this is actually enforced - **The write path.** A shared client wrapper that every team uses, which refuses or trims a write past the bound. This is the only place where enforcement is genuinely enforcement rather than intention. - **Review.** The standing rule is that any shape which can grow without a natural bound must name its bound before it ships — which forces someone to do the multiplication while it is still cheap. - **Measurement.** Knowing which entries are approaching the bound on a live tier, without hurting the traffic while you find out, is an operating practice with its own techniques. The rule's contribution is a number for that measurement to compare against. - **Ownership.** The key convention should encode which team owns an entry, so a breach arrives with a name attached rather than as a prefix nobody recognises. ## The exit clause The most valuable thing to say here is when the answer is that the data should not be in this keyspace at all. An append-only sequence that must keep every member, a record whose retention is set by a business rule, anything an auditor will ask about in two years: these are durability requirements wearing an in-memory costume, and no bound you agree to will survive contact with the requirement that nothing may be discarded. The move is to put the sequence in a store built for it and keep in the volatile tier only the bounded view that serving actually needs. A rule that cannot say this out loud ends up being quietly broken by the one team whose data genuinely could not be capped. ## What the rule is really buying Not tidiness. Three things: a tier whose worst operation is bounded and therefore predictable, a shared budget that one team cannot consume by accident, and a keyspace that is legible to whoever is on call in two years — because a convention that encodes owner, shape and bound turns an unfamiliar key into something a stranger can reason about at three in the morning. That last one is the durable payoff, and it is why the rule is written down rather than remembered.
- Why state a member bound as well as a byte bound?They fail differently. A collection can stay modest in bytes while holding millions of tiny members, which makes any whole-entry operation expensive in time even though the memory figure looks fine. And a few large members can breach a byte bound at a member count nobody thought unusual. Bounding one does not bound the other.
- Which part of this belongs to whoever operates the tier rather than to the rule itself?The measurement. Finding which entries are approaching the bound on a live tier, and doing it without hurting the traffic, is an operating practice with its own techniques and its own risks. The rule's job is to fix the number and name the owner, so that the measurement has something to compare against.
saying these in an interview costs you the question
- The store's own limit is the limit, design right up to it
- No rule is needed, teams will notice when it gets slow
- A review checklist is enough, the write path need not know the bound
- One byte bound covers it, member count does not matter
- The tier will evict the oversized entry, so it bounds itself