skip to content

Fragmentation & Resident Size

Memory freed inside the process is often not returned to the operating system, and under fixed-size classes it hides inside the store's own total. Telling either from a leak is the skill.

on this pageshow

questions

5

A store reports 6 GB held in entries while the operating system shows the process resident at 9 GB — what does each number measure?

level: middleimportance: must knowfreq 62%

answer

  1. two accountants, one process
  2. store counts entries, kernel counts pages
  3. rounding, unreturned blocks, connection buffers
  4. a small gap proves nothing alone
  5. ratio needs uptime and history

basics

~20 s

Data size is what the store attributes to its own entries; resident size is what the operating system says the whole process occupies. The gap between them holds allocator rounding, blocks the process freed but never gave back, and memory held for connections rather than for entries.

solid answer

~50 s

Two different accountants are reporting. The store's **data size** is its own sum over the entries it believes it holds — key bytes, value bytes and whatever metadata it keeps per entry. **Resident size** is the operating system's view of the pages the process actually occupies, and it is the figure that decides whether the machine starts paging to disk or reaches for its out-of-memory kill. The 3 GB between them is not automatically waste: it holds the rounding the allocator applied to every request, blocks freed inside the process and not handed back, and memory the process holds on behalf of connections and followers rather than on behalf of entries. A ratio a little above one is the ordinary state of a healthy process. What a particular ratio means is not answerable from the snapshot — it needs how long the process has been up and what it has been asked to store.

go deeper

for a junior

Know that there are two memory numbers, not one: what the store says it is holding in entries, and what the operating system says the whole process occupies. Quote the second when someone asks whether the machine is in trouble.

for a middle

Explain what lives between them — rounding the allocator applied to every request, blocks freed inside the process and never handed back, and buffers held for connections rather than entries — and say why a ratio slightly above one is ordinary.

for a senior

Refuse to judge the ratio from a snapshot. Ask for uptime and what the workload has been doing, and be explicit about which figure you compare against the machine limit and which against the store's own configured ceiling.

for a principal

Point out that where the waste gets counted is itself implementation-dependent: a store handing out fixed-size blocks counts the wasted remainder inside its own total, so a healthy-looking ratio is not evidence of a healthy process.

## Four numbers, all called "memory" When somebody says an in-memory store "is using 9 GB", at least four different figures could be meant, and they do not agree with one another: | Figure | Who reports it | What it counts | |---|---|---| | Payload bytes | Nobody — you compute it | The key and value bytes you actually wrote | | Data size | The store | What the store attributes to its entries: payload plus its own per-entry metadata | | Resident size | The operating system | The pages the whole process currently occupies | | The machine limit | The machine or container | The ceiling the process is not permitted to cross | The question above hands you the middle two, and the whole skill is knowing that they are different measurements rather than one measurement taken twice. **Data size** is bookkeeping the store does for itself. It walks its own structures and adds up what it thinks each entry costs. It is an honest number about entries and a silent number about everything else. **Resident size** is the operating system counting pages. It has no idea what an entry is. It knows only that this process has touched these pages and they are currently in memory. It is the number that matters when the machine decides who is using too much. ## What sits in the gap - **Allocator rounding.** A request for 37 bytes does not get 37 bytes. The allocator rounds every request up to the next step it supports, and the remainder is real memory that no entry is using. Multiplied across tens of millions of entries, the rounding alone is a visible number. - **Freed blocks the process kept.** When an entry is deleted, its bytes go back to the process's own free lists, not to the operating system. The process keeps them ready for reuse, and its resident size does not fall. - **Memory held for connections and followers.** Buffers for clients that are reading slowly, and buffers holding a stream of changes for a follower, are memory the process occupies on behalf of connections rather than on behalf of entries. The store usually does not count them in its entry total. - **The process itself.** Code, thread stacks, internal structures and anything the store allocates that is not attributable to one entry. - **Transient duplication.** While a whole-keyspace copy is being written out in the background, pages changed during that window can exist twice, so resident size rises for the duration and falls again afterwards. ## Why a gap is the normal state A general-purpose allocator cannot give memory back to the operating system at the granularity at which a store frees it. The operating system deals in pages and larger spans; the store deals in one entry at a time, of many different sizes, scattered wherever there was room. A page that still holds one live entry cannot be handed back no matter how empty the rest of it is. So a ratio slightly above one is not a symptom. It is what a long-running process that allocates in small pieces looks like. The ratio — data size against resident size — is worth tracking precisely because it is normally boring. A number that sits in a narrow band for weeks and then leaves it is telling you something happened; the number on its own, read once, is telling you almost nothing. ## Where implementations diverge This is the part that makes the two-number model something other than one store's trivia. - **Not every store reports both.** How much of its own accounting a store publishes, and what it calls the figures, differs widely. Some expose a detailed breakdown, some little more than a total, and some effectively leave you with the operating system's number and an entry count. - **Where the waste is counted differs.** A store that allocates from fixed-size blocks hands an entry the smallest block that fits and counts that **whole block** as used. The remainder is wasted, but it is wasted *inside* the store's own total, so the data-to-resident ratio looks healthy while a large share of the bytes is unusable. Under a general allocator the same class of waste tends to land *outside* the store's total, in the gap. A small gap therefore proves nothing on its own — you have to know which scheme you are looking at. - **Allocators differ in what they return.** Some return large empty spans to the operating system fairly eagerly; others hold on. This is not a property of in-memory stores at all, but it changes what you see. ## Reading the ratio in practice 1. **Say which figure you have.** Before answering "is 9 GB a problem", state whether that is the store's entry total, the process's resident size, or the limit it is running under. Three defensible answers contradict each other otherwise. 2. **Compare each against the right ceiling.** Resident size goes against the machine or container limit. The store's own total goes against the store's own configured ceiling. Crossing them over is how teams arrive at the out-of-memory kill with the store's dashboard still green. 3. **Get uptime and workload history before judging the gap.** The same ratio is unremarkable on a process that deleted millions of entries this morning and alarming on one whose entry count has been flat for a month.

  • Which of the two figures does the machine actually act on?
    Resident size. The operating system never reads the store's accounting; it counts the pages the process occupies, and that count is what drives paging to disk and, at the end, the operating system's out-of-memory kill. The store's own entry total is the right figure to compare against the store's own configured ceiling, and only that.
  • Can resident size ever be lower than the store's reported data size?
    Yes. Pages that were allocated but never touched, or that were touched long ago and have since been paged out to disk, are not resident even though the store still counts the bytes in them. A process under memory pressure can therefore report more entry data than the operating system says it holds — which is a sign of paging, not of better accounting.
  • Are you blind on a store that publishes no data-size figure of its own?
    Coarser, not blind. You still have resident size and an entry count, so you can track bytes per entry over time and watch the trend. What you lose is the ability to split the total into entries against everything else — which is exactly the split the gap question turns on, so you compensate with history instead of decomposition.

saying these in an interview costs you the question

  • Reads the store's own total as what the machine sees.
  • Calls any gap between the two figures a leak.
  • Assumes deleting entries shrinks the process immediately.
  • Expects the two numbers to agree on a healthy process.
  • Thinks a small gap proves no memory is being wasted.
open as a page

A store up 40 days shows resident size at twice its entry total, and the ratio has climbed every day this week — how do you tell unreturned memory from a leak?

level: seniorimportance: must knowfreq 58%

basics

~20 s

The snapshot cannot decide it; the trend and the workload history can. Establish which of the two figures is moving, what the workload did before the gap opened, whether the gap is flat or climbing, and whether a restart returns it and the climb then resumes on the same curve.

open as a page

A store had a third of its entries deleted an hour ago; its data total fell but the process never shrank — why?

level: middleimportance: should knowfreq 55%

basics

~20 s

Freeing an entry returns its bytes to the process's own allocator, not to the operating system. Those blocks wait in free lists for reuse, and the process shrinks only when a whole region happens to empty — which a deletion scattered across the keyspace rarely produces.

open as a page

A store allocating from fixed-size blocks had its small values purged and replaced with much larger ones; far fewer entries now fit in the same total — why?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Free space belongs to the block size it was cut for. Blocks released by the purged small values can only take values that fit them, so the larger values must come from the classes that serve their size — and the memory freed elsewhere is present, free, and unusable.

open as a page

A long-running in-memory store's resident size only ever grows between restarts, and only a restart returns it — how do you plan around that?

level: principalimportance: should knowfreq 35%

basics

~20 s

Treat the gap as a rate rather than a number: measure how fast it grows and how much room is left, then choose between changing the workload that produces it, letting the store consolidate free space where it can, and scheduling the restart as a planned availability event instead of an incident.

open as a page