skip to content

A bulk load briefly made every collection in a store large; the collections were trimmed back, yet the store's own data size stayed high — why?

level: seniorimportance: should knowfreq 46%

answer

  1. the peak size decided the layout
  2. trimming removes elements, not the structure
  3. checking on every removal would thrash
  4. attributed data size, not resident size
  5. rewrite the entry to re-decide it

basics

~20 s

Stores that swap to a general representation at the threshold usually never swap back when the value shrinks: re-checking on every removal costs work and would thrash at the boundary. The value keeps the larger layout until it is written again.

solid answer

~50 s

The load pushed every value past its **promotion threshold**, so the store laid each one out in the **general representation** — a slot and its bookkeeping per element. Trimming them removed elements, but in most stores that keep two layouts the change is effectively one-way: checking on every removal whether the value could return to the packed form would put work on the hot path for a rare benefit, and values hovering at the boundary would rebuild over and over. So the entries are small again and expensively laid out, and the store's own data size — what it attributes to its entries, rather than what the operating system sees the process holding — reflects that. The fix is to write the affected entries again from scratch, so the layout is chosen at their current size. Stores holding opaque bytes show none of this.

go deeper

for a junior

Remember that the internal layout of a value is chosen when it is written, based on its size at that moment, and that making it smaller later does not automatically undo that choice.

for a middle

Explain why a store skips the reverse check: it would cost work on every removal and would rebuild repeatedly for values sitting near the boundary, with no visible benefit to the caller.

for a senior

Separate the two diagnoses out loud. High attributed data size with small values points at the layout; high resident size with normal attributed size points somewhere else entirely, and rewriting one entry settles which you have.

for a principal

Treat it as a migration design problem: the peak size during a backfill sets the layout for the whole keyspace, so the import plan, not the steady state, is what needs reviewing.

## What the numbers are actually saying Three facts together point at one cause: the entry count is unchanged, the values are genuinely small again, and the figure that stayed high is the store's **data size** — what the store attributes to its own entries — rather than only its **resident size**, which is what the operating system reports the process holding. That last distinction is what rules out the competing diagnosis. If only resident size were high, the suspect would be memory the process is holding that belongs to no entry. Here the store itself still believes these entries are expensive, and it is right: they are. ## Why the layout does not come back on its own When the bulk load pushed each value past its threshold, the store laid it out again in the general representation, giving every element its own place in a lookup structure. Trimming the values removed elements from that structure. It did not reconsider the structure. In most stores that keep two layouts, the promotion is effectively one-way, for reasons that are all about the hot path: - **Checking costs something on every removal.** The decision to pack requires knowing the current element count and the largest element size. Recomputing that on every removal spends time on the common path to buy a saving that almost never applies. - **It would thrash.** A value oscillating around the boundary would be laid out again in both directions repeatedly, and each rebuild is proportional to the number of elements. The pathological workload would be an ordinary one. - **The rebuild is not free when it does apply.** Going back means walking every element and writing a new contiguous run, so the store would be paying a rebuild inside a removal the caller expects to be trivial. - **Nothing in the caller's experience improves.** The results are identical either way. The only beneficiary is memory, and memory has no way to ask. ## Reading the evidence | observation | what it suggests | what distinguishes it | |---|---|---| | data size high, values small, entry count flat | values still in the general layout | rewriting one entry and re-measuring moves the figure | | resident size high, data size normal | memory held by the process but owned by no entry | the store's own attribution is unremarkable | | data size high, values genuinely large | the values really are that big | inspecting an entry's element count settles it immediately | | data size climbing with a flat workload | entries being added or growing somewhere | the entry count or the size distribution is moving | ## What gets the memory back 1. **Write the entry again from scratch.** Build the value fresh — into a new entry and swap, or delete and write — so the store picks the layout for the size it has now. This is the direct, targeted remedy and it is measurable one entry at a time. 2. **Reload the keyspace.** Anything that rebuilds every value from a copy re-decides every layout at current sizes, because the layout is chosen at write time. That is a heavy instrument for this problem, and on a tier holding state with no source of truth it is not an instrument you reach for casually. 3. **Not the memory ceiling.** Pressure removal takes whole entries away; it does not make a surviving entry cheaper. If you are counting on it to clean this up, you are counting on losing data to fix an accounting problem. ## The shape of the trap The specific story matters less than its shape: **something temporarily inflated every value, and the inflation left a permanent mark.** A backfill that writes a whole history before trimming to a window. A migration that lands both the old and the new fields before dropping the old. An import that builds a collection element by element from a source that was not sorted, so the working value is large before it is pruned. In each case the peak size, not the final size, decided the layout of the entire keyspace. The defence is to write values in the shape you intend to keep. Where the import genuinely must inflate them, plan a rewrite pass afterwards, and measure what the store attributes to the affected entries before declaring the migration finished — because the workload has by then gone quiet and nothing else will raise the alarm. ## Where stores differ - Stores that hold values as opaque bytes have one layout and nothing to demote; every write replaces the value wholesale anyway. - Stores that keep two layouts choose their own axes and thresholds, so which values were promoted by the same load is not portable knowledge. - A few designs re-evaluate the layout on paths that already rebuild the value, so a wholesale overwrite gets it back while an element-by-element trim does not. - Stores that allocate from fixed-size blocks have an analogous but separate effect: a value that shrank does not move itself into a smaller block. Assert none of these as the model. The portable claim is narrower and still useful: **the layout is decided at write time, from the size at that moment, and shrinking a value later is not a write of the whole value.**

  • Why not just re-check the layout whenever an element is removed?
    Because it puts work on the hot path for a benefit that almost never applies, and a value oscillating around the boundary would be laid out again in both directions repeatedly. Each rebuild costs time proportional to the element count, so the cheap common case would pay for the rare expensive one.
  • How do you confirm it is the layout rather than memory the process is holding for no entry?
    Look at which figure is high. If the store's own attribution to its entries is high while the values are small, the entries themselves are expensive. Then rewrite one affected entry from scratch and measure again: if that entry's attributed cost drops sharply, the layout was the cause.
  • How would you keep a future backfill from doing this?
    Write values in their final shape rather than inflating and trimming, and stage the import so the working value never exceeds what you intend to keep. Where that is impossible, schedule a rewrite pass over the affected entries and check the attributed size before you call the migration done.

saying these in an interview costs you the question

  • Assumes a value returns to the packed layout as soon as it shrinks.
  • Blames the allocator when the store's own attributed size is what stayed high.
  • Thinks removing elements always returns the memory those elements used.
  • Expects the memory ceiling to fix it by removing entries under pressure.
  • Believes every store rebuilds its value layouts in a background pass.