A per-request arena is reset on every response, yet the service's peak memory far exceeds its live data. Why?
answer
- held, not lost
- cumulative, not live
- one release point per phase
- multiply by phases in flight
- shorten the phase before anything else
basics
~20 sBecause an arena holds everything allocated during the phase, not what is still in use. Intermediate scratch that died early is still occupying the region until the reset, and that total is multiplied by the number of requests in flight.
solid answer
~50 sAn arena has exactly one release point, so its high-water mark tracks **total bytes allocated during the phase**, while a general-purpose allocator's footprint tracks roughly the **live set**. If a request builds 200 scratch buffers of 1 MB each but never holds more than 3 at once, the arena holds 200 MB where per-object release would have held about 3 MB — and with 64 requests in flight that is 12.8 GB rather than a few hundred. A single slow or oversized request makes it worse, because its whole region is pinned for as long as the request runs. The levers are all about shortening the phase or bounding it: reset inner scopes at sub-boundaries such as one page or one section, stream output instead of building it entirely in memory, cap region growth and fail the request against a declared budget, and keep genuinely long-lived data out of the arena.
go deeper
Remember that an arena keeps everything the phase allocated until the reset, so memory in use is not the same as memory still needed.
Explain the swap: a general allocator's footprint tracks the live set, an arena's tracks total allocated per phase, and say what the churn ratio between them is.
Diagnose it as retention rather than a leak, and reach for the inner reset boundary, streaming, a per-request cap and a concurrency limit in that order.
Own the sizing formula — maximum bytes per phase times maximum phases in flight — and defend the cap policy, including what the service does to a request that exceeds it.
## Where the number comes from Two different quantities drive footprint, and the arena swaps one for the other. - Under **per-object release**, memory in use at any instant is roughly the **live set**: things released come back immediately and are available for the next allocation. - Under an **arena**, memory in use is the **cumulative total allocated since the last reset**, because there is no mechanism to take anything back early. Garbage produced in the middle of the phase is indistinguishable from live data until the boundary arrives. So the design's defining advantage — no per-object release — is exactly the cause of its footprint profile. This is not a leak. Nothing is lost, nothing is unreachable, and the reset returns all of it. It is **retention by design**, and it has to be budgeted rather than debugged. ## A worked count Take a report renderer with a per-request arena: 1. The request renders 200 sections, and each builds a 1 MB scratch buffer, formats it, appends the result, and moves on. 2. At any instant at most 3 of those buffers matter — the one being built, the one being formatted, the output accumulator. 3. With per-object release the footprint sits near **3 MB**. With an arena it sits at **200 MB** at the end of the request, because all 200 buffers are still in the region. 4. At 64 requests in flight, that is 64 x 200 MB = **12.8 GB** rather than a few hundred megabytes. The ratio is the phase's allocation churn: the more the phase allocates and discards internally, the worse the arena's profile, and the churn ratio is the single number to measure before adopting the design. ## The long phase The second effect is time, not volume. Because the reset is the only release, **the region is pinned for the whole duration of the phase**. One request that stalls on a slow dependency holds its entire region while it waits, doing nothing with it. A design that assumed phases are short and uniform develops a fat tail exactly where its latency has one, and the two failures correlate: the slowest requests are usually the ones that also allocated the most. ## Symptom, cause, lever | Symptom | Cause | Lever | |---|---|---| | Footprint far above live data | Internal churn retained until reset | Sub-arenas reset at inner boundaries | | Footprint scales with concurrency | One region per in-flight phase | Bound concurrency, or bound the region | | A few requests dominate the peak | Oversized or slow phases pin whole regions | Per-request cap, then fail explicitly | | Peak grows with input size | Whole result built before it is emitted | Stream output as it is produced | ## The four levers, in the order to try them 1. **Shorten the phase.** If the work has an inner boundary — one page, one section, one batch of rows — give it its own arena and reset there. This is usually the largest win and costs nothing but discipline, because it turns one 200 MB region into a 1 MB region reused 200 times. 2. **Stream instead of accumulating.** If the output can be emitted progressively, the accumulator stops being the dominant allocation and the phase's total falls with it. 3. **Cap and fail.** Give the region a declared maximum and make exceeding it a clean, attributable failure of that one request. An uncapped arena turns an unusually large input into a process-wide memory event; a capped one turns it into one bad response with a name on it. 4. **Move long-lived data out.** Anything that must survive the reset is being copied out anyway. Allocating it in the arena first means it is held twice at the boundary. ## What this means for capacity planning The sizing formula for an arena-based service is not the live set. It is: **peak footprint ≈ (maximum bytes allocated in one phase) x (maximum phases in flight)** Both factors are things the service controls — the first by the four levers above, the second by its concurrency limit — which is the redeeming feature of the design. The footprint is high, but it is **predictable and bounded by numbers you set**, rather than emerging from an allocator's history. Teams that adopt arenas and then discover their footprint usually skipped the first factor; teams that succeed with them measured per-phase allocation first and chose the concurrency limit second. One last property worth stating plainly: the reset reclaims bytes and runs no per-object cleanup. If the phase also opened handles or took locks, shortening the phase for footprint reasons does not release those, and they need their own explicit discipline.
- Is this a memory leak?No. Nothing is unreachable and nothing is lost: every byte comes back at the reset. It is retention by design, so the right response is to budget it — shorten the phase, cap the region, bound concurrency — rather than to hunt for a missing release that does not exist.
- Which single change usually cuts the peak the most?Finding an inner boundary and resetting there. Work with sections, pages, rows or batches usually has one, and moving the reset inward turns one region sized to the whole phase into a small region reused many times. It costs only the discipline of not letting anything escape the inner scope.
- How should the service behave when one request's arena exceeds its budget?Fail that request cleanly and attributably, at the cap. An uncapped region lets one oversized input become a process-wide memory event affecting every other request in flight; a cap converts it into a single bad response you can name, count and alert on.
saying these in an interview costs you the question
- Calls the retained scratch memory a leak and hunts for a missing release.
- Sizes the service from its live data rather than per-phase allocation times concurrency.
- Assumes unused space inside the region is reusable before the reset.
- Ignores that one slow request pins its whole region while it waits.
- Leaves the region uncapped, so one large input becomes a process-wide event.