A worker is killed outright while the engine's own numbers show its budget well below full. What explains that?
answer
- three numbers, one word
- the engine reports only what it counts
- killed, not refused an allocation
- the gap is runtime plus user allocations
- raising the budget shrinks the uncounted room
basics
~20 sThe engine reports only the memory it accounts for. The platform's ceiling covers the whole process, including the language runtime, buffers and user-function allocations the engine never counted - so the process crosses a line the engine could not see.
solid answer
~50 sThree different numbers answer to the word memory here, and only one of them is in the engine's report. There is **the engine's accounted budget**, which it divides into regions and tracks; **the process's actual footprint**, which also includes off-budget overhead - the language runtime itself, native and network buffers, and anything a user-supplied function allocates outside the engine; and **the limit the platform enforces**, a ceiling applied to the whole process by the platform rather than the engine. The platform enforces it by killing the process outright rather than by making one allocation fail, so there is no engine error and no stack trace: a unit of work simply disappears and is retried elsewhere, often repeatedly, until the job gives up. The tell is the *absence* of an engine-side memory error. An allocation the engine refused would have been reported, attributed and often degraded to spilling instead.
go deeper
Recall that the engine's report covers only the memory it manages, and that a worker can be stopped by the platform for the whole process's size even when that report looks fine.
Explain what sits in the gap - the language runtime, native and network buffers, and whatever user-supplied code allocates - and why the failure arrives with no engine-side error.
Show the diagnosis in order: name which of the three numbers moved, read the absence of an engine error as evidence, and resist enlarging the accounted budget as a reflex.
Your angle is the default that prevents it at scale: how much room the platform leaves beneath the ceiling for uncounted bytes, and what review a team owes code that allocates per record.
## Three numbers, and the one that killed you - **The engine's accounted budget.** What the engine hands out to operators and to kept results, and what its per-unit report describes. It is a real number and it is not the whole story. - **The process's actual footprint.** What the machine sees the worker process holding. It is the accounted budget plus **off-budget overhead**: memory the process really uses that the engine's accounting never counted. - **The limit the platform enforces.** A ceiling the platform, not the engine, applies to the whole process, and enforces by killing it outright rather than by refusing one allocation. The scenario in the question is the gap between the first and the second being charged against the third. The engine was telling the truth about the region it manages, and the region it manages was never the thing being measured at the ceiling. ## Why the kill looks nothing like an allocation failure | | allocation the engine refused | the platform killing the process | |---|---|---| | who noticed | the engine's own allocator | the platform, outside the engine | | what you see | an error attributed to a unit of work | a worker that stops existing, with no engine message | | usual first response | spill to disk attached to the worker, or fail that unit | the unit is rescheduled on another worker | | does the engine's report show it | yes, in that run's numbers | no - the last numbers it wrote predate the kill | | retry behaviour | may well succeed after a replan or a spill | reproduces, because the same footprint recurs | The abruptness is diagnostic. A killed process cannot flush its own final numbers, which is why the evidence is a gap in the record rather than an entry in it. ## What lives in the gap - **The language runtime itself** - its own structures, thread stacks and internal buffers, which exist before your first record. - **Native and network buffers** allocated outside the engine's accounted regions for reading, writing and moving bytes between workers. - **User-supplied function allocations.** The largest and most common cause: a lookup table loaded once per unit of work, a parser that buffers a whole document, a native library with its own allocator, or records held in a closure that never lets go. The engine sees the function as a black box and cannot account for what it allocates. - **Concurrency multiplying all of the above.** Several units of work run at once inside one worker, and each of them may carry its own copy of the uncounted allocations, so the gap is roughly per unit rather than per worker. ## What actually helps 1. **Decide which number moved before changing anything.** If the engine's report shows the accounted budget nowhere near full, the accounted budget is not your problem and enlarging it is not your fix. 2. **Leave the process room the engine will never ask for.** The gap has to fit under the ceiling alongside the accounted budget. Enlarging the accounted budget without enlarging the ceiling makes the situation strictly worse, because the uncounted region now has less room. 3. **Look inside the user-supplied code first.** Anything allocated per record, per unit of work, or loaded eagerly at start-up is uncounted by construction. A per-record allocation that is not released is the classic finding. 4. **Count the concurrent units.** If each concurrent unit carries an uncounted buffer, the footprint scales with concurrency even though the accounted budget does not visibly move. 5. **Watch for repeated kills of the same shape.** A single kill can be a noisy neighbour on the machine; the same unit dying on three different workers is a property of the work, not of a machine. ## What varies between engines This is one of the places where engines of this class diverge most. Some ask you to declare an allowance for the uncounted region explicitly alongside the accounted budget, so the two are sized together; others expect the operator simply to leave room, and will happily let you configure a budget that cannot fit under the ceiling. The reported peak also means different things: usually it is the peak of what the engine accounted for, not the process's resident footprint, so two engines' peak figures are not comparable and neither is a statement about the ceiling. And on engines holding records as ordinary language objects rather than as packed bytes the engine manages, the same records occupy several times more, which widens the gap for identical data. ## The line to hold in an interview Name the three numbers, say which one was exceeded, and say why the evidence for it is the *absence* of an engine-side error. A candidate who reaches for a bigger budget before saying which number moved has answered a different question - and has usually made the real one worse. What the ceiling is, who granted it and how the platform applies it is the platform's subject, not the engine's; what the engine owes you is an honest account of the region it manages and room left for the region it does not.
- How do you tell this apart from an allocation the engine itself refused?By whether anything was reported. A refused allocation is the engine's own event: it is attributed to a unit of work, appears in that run's numbers, and often degrades to writing part of the working set to disk attached to the worker instead of failing. A kill leaves no engine message at all, because the process was removed rather than told no.
- Why does retrying the unit of work usually not help?Because the footprint that crossed the ceiling is a property of the work and the code, not of the machine that happened to run it. The same slice of data and the same user-supplied function allocate the same uncounted bytes on the next worker. Repeated kills of the same unit on different workers are the confirmation.
- Why can enlarging the engine's accounted budget make this worse?The ceiling counts the whole process. Giving the engine a larger accounted region does not raise that ceiling, so the uncounted region - runtime, buffers, user allocations - has less room beneath it than before. The process then reaches the ceiling sooner, while the engine's report looks even more comfortable.
saying these in an interview costs you the question
- Assumes a killed worker means the engine's accounted budget was full
- Raises the accounted budget without raising the process ceiling
- Expects an engine error message for a process the platform killed
- Blames the machine when the same unit dies on several workers
- Ignores what a user-supplied function allocates outside the engine
- Uses one word for the budget, the footprint and the ceiling