skip to content

A dashboard for an in-memory store shows a single 'memory used' figure - which four quantities could it be, and what does each mean?

level: juniorimportance: must knowfreq 72%

answer

  1. one label, four different numbers
  2. server's count versus the operating system's
  3. the ceiling is a policy, not a measurement
  4. ceiling below host limit decides who acts
  5. which number is closest decides the failure

basics

~20 s

Four quantities hide behind one 'memory used' figure: stored-data size (the server's count for its entries), resident footprint (what the operating system sees), the memory ceiling the server enforces, and the host or container limit above it.

solid answer

~50 s

There is no single memory number for this component. **Stored-data size** is what the server attributes to the entries it currently holds, computed by the server itself. **Resident footprint** is what the operating system sees the process holding, which includes everything the server's own accounting does not attribute to entries. **The memory ceiling** is the bound the server enforces on itself, on stores that have one - it is a configured policy number, not a measurement, and it does not move when memory does. Above it sits **the host or container limit**, enforced against the resident footprint, where the process is killed rather than asked to behave differently. Saying 'memory is at ninety percent' without naming which of the four is the reading hides the whole diagnosis, because the gap between two of them is usually the finding.

go deeper

for a junior

Recall that 'memory used' is four different numbers and say which one you are quoting before you say anything else about it. Knowing that the server's count and the operating system's count are separate figures is most of the credit here.

for a middle

Explain who computes each number and what each one excludes, and pair each one with the failure it predicts: the server changing behaviour at its ceiling, or the process being killed at the host limit.

for a senior

Show that you alert on both ends and on a sustained window, and that you check where the ceiling sits relative to the host limit before trusting any of these graphs. Say what a flat line at the ceiling actually means.

for a principal

Frame it as who gets to act first when the tier runs out of room - the server, with a behaviour change teams can see, or the kernel, with a silent kill - and make that a standing configuration contract rather than a per-instance accident.

## Four numbers, one label An in-memory store does not have *a* memory number. It has at least four, they are computed by different parties, they measure different things, and the failure each one predicts is different. A dashboard tile labelled `memory used` may be any of them, and a statement like "the tier is at ninety percent" is ungradeable until the quantity is named. | Quantity | Who computes it | What it is | What it silently excludes | |---|---|---|---| | **Stored-data size** | the server, from its own accounting | the total the server attributes to the entries it currently holds | overhead the server does not attribute to entries - allocator bookkeeping, connection and reply buffers, work in progress. Exactly which of these sit inside the figure differs by store | | **Resident footprint** | the operating system | the physical pages the process is holding right now | any explanation of *why* it is that size; it cannot separate entries from overhead | | **The memory ceiling** | configured; the server enforces it on itself | the bound at which the server changes behaviour - removing entries to make room, or refusing the write | it is a policy number, not an observation. It does not move when memory moves, and not every store in this class has one | | **The host or container limit** | the kernel, container runtime or orchestrator | the point at which the process is killed outright | it is enforced against the resident footprint, never against the server's own accounting | ## Why the label matters more than the value The four numbers do not move together, and each pair that diverges is a different finding: - **Stored-data size against the ceiling** decides when the *server* changes behaviour. This is the pair that predicts entries starting to disappear, or writes starting to be refused. - **Resident footprint against the host limit** decides when the *kernel or orchestrator* changes behaviour. This pair predicts a process that is killed with no warning and comes back holding nothing. - **Resident footprint against stored-data size** is memory the process holds but does not currently attribute to your entries. The gap is real and it is normal for it to be non-zero; what causes it, and what to do about it, is a separate subject from reading the signal. - **The ceiling against the host limit** is the most consequential configuration decision behind these graphs. If the ceiling sits comfortably below the host limit, the server gets to act first and you get a behaviour change you can see. If the ceiling is absent, or set at or above the host limit, the kernel acts first and you get a kill. ## What each reading alone licenses you to conclude 1. **Stored-data size pinned just under the ceiling, flat.** The server is holding at its bound. Flatness here is the *expected* shape, not evidence of health - it usually means the server is making room continuously. 2. **Stored-data size climbing steadily with no ceiling in sight.** Nothing is bounding growth inside the server, so the host limit is the only thing that will stop it. 3. **Resident footprint well above stored-data size.** The server believes it has room the operating system does not agree exists. Any ceiling derived from the server's own figure does not bound what the host sees. 4. **Resident footprint dropping to near zero and climbing again.** The process restarted. Everything the tier held is gone unless something restored it. ## Where stores in this class differ - **Not every store enforces a ceiling on itself.** Some take a fixed capacity at start-up, some are bounded only by the host, and among those that have a ceiling, the behaviour at it varies - removing entries to make room on some, refusing the write on others. - **What counts against the ceiling varies.** Whether reply buffers, connection state or the extra memory a background whole copy costs are inside the server's figure or outside it is a per-store property, so two stores reporting the same number are not reporting the same thing. - **Managed deployments often expose only a subset.** Where you cannot reach the host, the resident footprint and the host limit may be visible only as a vendor-shaped percentage, and you must confirm which quantity it is derived from before writing an alert against it. ## Turning it into an alert Alert on **both ends**, because they fail differently and one going quiet does not make the other safe: stored-data size approaching the ceiling (the server is about to change behaviour), and resident footprint approaching the host or container limit (the process is about to be killed). Write both as prose arithmetic against a sustained window - for example, resident footprint within a tenth of the host limit, sustained over ten minutes - rather than on a single scrape, which catches a transient spike and pages someone for nothing.

  • Stored-data size sits comfortably under the ceiling while the resident footprint sits above it. What does that fact alone let you conclude?
    That the server believes it has room the operating system does not agree exists, so a ceiling derived from the server's figure is not bounding what the host sees. The limit that will actually fire is the host or container one, and it fires as a kill rather than a behaviour change. Naming the cause of the gap is a separate investigation; the signal reading stops here.
  • Why is a memory ceiling set at or above the host limit worse than having no ceiling at all?
    Because it looks like protection and provides none. The server never reaches its own bound, so it never removes entries or refuses a write, and the first thing that happens is the kernel killing the process - losing everything the tier held, with no graded warning in between. A ceiling is only useful when it sits far enough below the host limit that the server acts first.
  • On a managed instance where you cannot reach the host, which of the four numbers can you still act on?
    Usually the server's stored-data size against the configured ceiling, plus whatever aggregate percentage the operator interface publishes. Before alerting on that percentage, confirm which quantity it is derived from - a figure based on the server's own accounting and one based on the process footprint predict different failures, and vendors differ in which they expose.

A delivery van has four weight numbers: the goods in the back, the loaded van on the weighbridge, the limit printed on the van's own plate, and the limit of the bridge ahead. They are all 'weight', they never match, and which one you are closest to decides whether the driver slows down or the bridge decides for you.

saying these in an interview costs you the question

  • Says 'memory is at ninety percent' without naming which of the four numbers
  • Treats stored-data size and resident footprint as the same figure
  • Assumes the memory ceiling is the same thing as the container limit
  • Believes every store in this class enforces a ceiling on itself
  • Reads a flat memory graph as proof the tier is healthy
  • Thinks the ceiling moves in response to actual memory use