A single-process data job 'ran out of memory' overnight; which distinct failures does that phrase cover, and how do you tell them apart?
answer
- four failures wear one name
- refused allocation against a process ended
- a full disk is a capacity symptom too
- peak during a step, not the average
- name the step, then the resource
basics
~20 sThe phrase covers at least four failures: a refused allocation inside the program, the process killed from outside for total resident size, a local disk filling with intermediates, and paging that ruined the wall clock. Different evidence, different step.
solid answer
~50 s'Out of memory' is a report, not a diagnosis. Four different exhaustions arrive under that one sentence. First, an allocation the program asked for was denied, so an error is raised inside the program and it names the step that asked. Second, nothing is raised at all: the process is ended from outside because its resident bytes crossed a ceiling at one instant, and the exit status is set by a signal. Third, a device-full error, because the run was writing intermediates to this machine's disk and the disk ran out before the work did. Fourth, nothing failed and the run simply took a working day, because pages were moving between memory and disk the whole time. Telling them apart means one thing: record peak resident bytes and elapsed time per step, not for the run as a whole.
go deeper
Recall that 'out of memory' is a report rather than a cause, and that a run can also die because the local disk filled or crawl because pages were moving to and from disk.
Explain the mechanics: what a refused allocation gives you that a process ended from outside does not, and why peak bytes during one step is the figure to look at rather than an average.
Show the diagnosis habit: per-step timings and per-step peaks collected before anything is changed, and a written sentence naming the resource and the step it was exhausted at.
Frame what the measurement is worth to the organisation: every later conversation about money or engineering time is unarguable without it, and guesses about which resource ran out are usually wrong and always expensive.
## The phrase hides four different failures A single-process data job that stops working is reported in one sentence — *it ran out of memory* — and that sentence is an umbrella. Underneath it sit at least four distinct exhaustions. They have different evidence, they happen at different steps, and they are not fixed by the same thing. - **A refused allocation.** The program asked for a block of bytes and the request was denied. An error is raised **inside** the program, and whatever the runtime prints points at the step that asked and often at the size it asked for. This is the most informative failure you can get. - **A process ended from outside.** No error reaches the program. Something outside it — the operating system under pressure, or a limit imposed on the process — ended it because the bytes attributed to that process crossed a ceiling at one instant. The exit status is set by a signal and the program's own logs simply stop mid-sentence. - **A local disk that filled.** The run was writing intermediates to this machine's own disk and reading them back, and the disk ran out before the work did. This is the same capacity failure of one machine, surfacing on a different device. - **Paging that ruined the wall clock.** Nothing failed at all. The machine kept the job alive by moving pages between memory and disk, and a run that should take twenty minutes takes nine hours. ## Reading the symptom backwards | what was observed | what it most directly says | what it rules out | |---|---|---| | Error inside the program naming a failed allocation | one request, of a known size, at a known step | the failure being invisible to the program — you have a name and a place | | Process gone, no error, exit status set by a signal | total bytes attributed to the process crossed a ceiling at one instant | a single named allocation being solely at fault; the peak was the sum | | Device-full error while a step was running | intermediates outgrew this machine's disk | main memory as the resource that was exhausted | | Run finished, hours late, storage busy throughout | pages moving between memory and disk | any hard ceiling being reached at all | Notice what the second row does **not** say. A process ended from outside does not prove that the dataset is larger than the machine. It proves that at one moment the process held more than it was allowed, and that moment may have been a duplicate of one column held briefly while a single step ran. ## The step matters more than the job A number for the whole run is nearly useless here, because exhaustion happens at a step. What you want in hand is small and cheap to collect: 1. **Elapsed time per step**, so the run is a sequence rather than a single duration. 2. **Peak resident bytes during each step**, not the average over the run and not the figure at the end — the failure is at the worst moment of one step. 3. **Free space on the device holding intermediates**, sampled while the run is going, not afterwards when the files have been cleaned up. 4. **The exit path itself**: an error inside the program, or an exit status set by a signal with nothing logged. These are different failures and the distinction is free to record. With those four you can write the sentence that actually means something: *the third step, the one that combined the two largest columns, held 41 GB on a 32 GB machine and was ended from outside.* Without them you have 'it ran out of memory', which names neither a resource nor a place. ## What varies, and why the same job fails differently on two machines Three things here genuinely differ by design, and assuming one of them is the class is how diagnosis goes wrong: - **Whether the program sees a refusal or is simply ended** depends on the platform and its policy for handing out memory it has not yet backed. On one configuration the program catches a clean failure; on another the same run disappears without a word. - **Whether the resident figure falls back after a peak.** Large allocations served their own mapping can be handed back when freed, so the number drops; allocations served from pooled arenas are usually retained for reuse, so the figure stays at its high-water mark with almost nothing live. Read a resident figure as *the worst this process reached*, not as *what it holds now*. - **Whether intermediates ever touch the disk.** Some execution models hold everything and fail rather than write; others write intermediates to local storage by design. The same data and the same logic therefore produce a memory failure on one and a disk-space or wall-clock failure on the other. ## What the record is for The point of all this is a named resource at a named step. That sentence is the only thing that makes the next conversation — whichever direction it goes — an engineering discussion instead of an argument between two guesses.
- The run finishes, but four times slower than usual, and storage is busy for the whole run. Which resource is that?Effective memory. Nothing hit a hard ceiling, so nothing raised an error; the machine kept the job alive by moving pages between memory and disk, and each of those moves costs orders of magnitude more than a memory access. The symptom is a long wall clock with low processor utilisation and constant storage activity.
- Why is peak bytes during a single step more informative than the average across the run?Because the failure happens at the worst instant, not at the typical one. A run can sit comfortably for an hour and die in the two seconds where one step held both its input and its output at once. An average over the run smooths away exactly the moment that mattered.
- The process was ended from outside. Does that tell you the data does not fit on this machine?No. It tells you that at one moment the process held more than it was allowed. That moment may have been a step holding a second copy of one column, or a structure that grew with the number of distinct values rather than with the input. Which of those it was is a different measurement.
saying these in an interview costs you the question
- Out of memory means the dataset is bigger than installed memory
- A process ended for memory always raises a catchable error first
- Steady usage over the run is the number that matters
- A device-full error is a storage problem, unrelated to the job's size
- The run finished, so memory was never the constraint
- One total duration is enough to identify the bottleneck