skip to content

A store's memory ceiling is set to the container's full 16 GB limit — why is that fatal, and what belongs in the gap?

level: seniorimportance: must knowfreq 58%

answer

  1. two limits, two different enforcers
  2. the process holds more than its entries
  3. followers, slow readers, a running copy
  4. polite refusal before the out-of-memory kill

basics

~20 s

The process holds memory the ceiling does not account for: buffers for followers and for slow connections, and duplication while a whole-keyspace background copy runs. With no gap, the machine limit is reached before the store's own ceiling.

solid answer

~50 s

A store's **memory ceiling** is a number about the store's own accounting; the **machine limit** is a number about the process. They are not the same, because the process also holds memory on behalf of things that are not entries. Three named consumers live in the gap: replication buffers holding writes for a follower, client output buffers holding replies a slow reader has not drained, and the transient duplication of pages modified while a whole-keyspace background copy is being written. They are not the only residents, and stores differ in which of them fall inside the ceiling's line. Set the ceiling at the machine limit and the store never reaches its own ceiling first; the machine limit arrives first instead, which means paging to disk and then the operating system's out-of-memory kill. The ceiling has to be reachable before the machine limit is.

go deeper

for a junior

Remember that the store's configured memory ceiling and the machine's memory limit are two different numbers, and that the first must be lower. The process holds more than just the entries you stored.

for a middle

Explain what the ceiling actually governs — the store's own accounting — and name at least one thing the process holds outside it. Say what happens when the machine limit is reached instead: paging first, then the process is killed.

for a senior

Defend the gap with named consumers rather than a percentage: buffers for followers, buffers for slow or large-reply connections, and duplication while a whole-keyspace copy runs. Alert on both numbers, because the store's accounting cannot see the buffers.

for a principal

Argue it as a choice of who enforces the limit. A reachable ceiling keeps the failure inside the store, where it is observable, configurable and survivable; a ceiling at the machine limit hands enforcement to the operating system, where the only outcome is total loss of the tier.

## Two limits, enforced by two different parties An in-memory store that offers a **memory ceiling** is offering you a limit *it* enforces, against *its own* accounting of what it is holding. The **machine limit** — the container's memory limit, or the physical memory of the host — is enforced by the operating system, against everything the process holds. That difference is the whole subject. The store's ceiling is a number it can act on politely and in advance. The machine limit is a number enforced by failure: the process slows to a crawl as memory is paged to disk, and then the operating system's out-of-memory kill ends it. Setting the first equal to the second means the store can never act first. ## What lives in the gap The process holds memory on behalf of things that are not entries. Three consumers are worth naming because they are the ones that surprise teams, and because each one is large, bursty, and driven by something other than how much data you stored: 1. **Replication buffers.** Where the tier has followers, the primary holds writes on their behalf — enough to feed a follower that is behind, and enough to carry one through a re-connection. The size is driven by write rate and by how far behind a follower is allowed to fall, not by the size of the dataset. 2. **Client output buffers.** A server can produce a reply faster than a connection drains it. The undelivered bytes queue per connection, so a slow reader, a stalled network path, or a caller that asked for a very large reply parks memory on the server. This grows with the number of connections and with the size of the largest reply, and it is a common reason a process that had "plenty of headroom" died. 3. **Transient duplication during a whole-keyspace background copy.** On stores that produce a point-in-time copy by writing out the whole keyspace while still serving traffic, pages modified during the copy are duplicated, so the process can hold meaningfully more than the data size for the duration. The extra is proportional to how much is written while the copy runs, not to the dataset — a quiet tier duplicates almost nothing and a write-heavy one can duplicate a large fraction. | Consumer | Held on behalf of | Applies when | What makes it grow | |---|---|---|---| | Replication buffers | Followers | The tier has replicas | Write rate; how far a follower may lag | | Client output buffers | Connections | Always, but sharply with slow readers | Connection count; largest reply size | | Copy-time duplication | The background copy | The store writes whole-keyspace copies | Write rate while the copy is running | These three are not an exhaustive list. What the allocator holds beyond the store's own accounting also sits in the gap, and so does anything else sharing the box. That is an argument for the gap being larger, never smaller. ## Why the ceiling at the machine limit inverts the failure With a gap, growth in stored data reaches the store's ceiling first, and the store gets to apply whatever posture it was configured with — the three possible outcomes at the ceiling are a subject of their own, and the point here is only that the store *has* an outcome to apply, in its own process, on its own terms, with its own logging. With no gap, none of that happens. The buffers and the copy push the process past the machine limit while the store's accounting still reads "under the ceiling", and the failure is delivered by the operating system: first the slowdown as memory is paged to disk — which looks to callers like the tier suddenly became a slow disk — and then the kill, which takes everything the tier held with it. The sentence to be able to say out loud: *a ceiling set at the machine limit converts a memory problem the store could have handled into a process death it cannot.* ## Where the machine limit is not the only shape - In a container with a hard memory limit, exceeding it kills the process, usually with no paging first — the slowdown warning you expect may simply not arrive. - On a shared host with paging configured, the process degrades for a long time before anything is killed, and the tier's latency is unrecognisable long before that. - Either way, the number worth alerting on is the process's resident size against the machine limit, *in addition to* the store's own accounting against its ceiling, precisely because the two can diverge. ## What varies between stores - Not every store offers a ceiling at all. Where it does not, the headroom argument moves into the node choice: you pick a machine whose limit leaves the same gap above what you intend to hold, and you watch it rather than configure it. - Stores differ in exactly which memory falls inside the ceiling's line — some count more of their own bookkeeping and buffers than others — so a gap sized by "what I know the setting excludes" is fragile, while a gap sized by named consumers and then verified against the process's resident size is not. - Not every deployment has all three consumers. No followers means no replication buffers; a store that never writes a whole-keyspace copy has no copy-time duplication. Headroom is defended item by item, not as a blanket percentage.

  • This deployment has no followers and never writes a whole-keyspace copy. Is headroom still needed?
    Yes, but less of it, and for different reasons. Client output buffers remain, and so does the gap between the store's own accounting and what the process actually holds. The difference is that a gap this small can be defended by measuring the resident size against the store's accounting under real traffic, rather than by reserving space for a burst you can name in advance.
  • Why does a slow client consume memory on the server at all?
    Because the server produces the reply before the connection has carried the previous one away. The undelivered bytes wait in a per-connection buffer, so a reader that stalls — or a caller that asked for a very large reply — parks server memory that has nothing to do with how many entries are stored, and several such callers at once can park a lot of it.
  • Which number should the capacity alert watch?
    Both, because they can diverge. One alert on the store's own accounting approaching its ceiling, with enough lead time to act, and a second on the process's resident size approaching the machine limit. The first catches data growth; the second catches the buffers and the copy, which is exactly what the first cannot see.

A lift's posted weight limit is set below the breaking strain of the cable, not at it. The posted number is the one the system can enforce politely — it declines the next passenger, in daylight, with a sign. The cable's number is enforced by the cable. Setting a store's memory ceiling at the machine limit is posting the cable's number: you have given up the polite refusal and kept only the failure.

saying these in an interview costs you the question

  • Sets the store's ceiling equal to the container memory limit
  • Assumes the ceiling accounts for everything the process holds
  • Believes reaching the ceiling and being killed are the same event
  • Reserves headroom as a round percentage with no named consumer
  • Applies identical headroom whether or not followers or copies exist