What actually happens when a process in a Docker container allocates past the container's `--memory` limit, and what does the `--memory-swap` flag change about that behavior?
answer
- limit = cgroup memory.max
- overcommit: killed on touch, not on malloc
- cgroup OOM kills inside the container only
- PID 1 killed => exit 137, OOMKilled=true
- --memory-swap = memory + swap TOTAL
basics
~20 sThe kernel first reclaims what it can; if that is not enough, the cgroup OOM killer SIGKILLs a process inside that container only. If the victim is PID 1 the container dies with exit code 137 and State.OOMKilled=true. --memory-swap sets the combined memory+swap ceiling; setting it equal to --memory disables swap.
solid answer
~50 s`--memory=512m` writes a hard limit into the container's cgroup. Allocation itself usually still succeeds — Linux overcommits — so enforcement happens when pages are actually touched. As charged memory nears the limit the kernel reclaims (page cache, clean pages, swap if available). If it cannot reclaim enough, the **cgroup** OOM killer SIGKILLs the highest-scoring process inside that container; nothing outside is affected. If that process is PID 1, the container stops with exit code 137 (128 + SIGKILL) and `docker inspect` reports `State.OOMKilled: true`, after which the restart policy applies. `--memory-swap` is the **total** of memory plus swap, not the swap size. Leave it unset and you get twice `--memory`; set it equal to `--memory` and swap is disabled so the container fails fast; `-1` allows unlimited swap. `--memory-reservation` is a soft target enforced only under host pressure. Note the app gets no chance to log — SIGKILL cannot be handled.
code
bash · 5 linesdocker run -d --name api --memory=512m --memory-swap=512m --memory-reservation=384m myapp:1.0
docker inspect -f '{{.State.ExitCode}} {{.State.OOMKilled}}' api
docker exec api cat /sys/fs/cgroup/memory.events
dmesg -T | grep -i 'memory cgroup out of memory'go deeper
Know that --memory is a hard cap, that exceeding it gets the process SIGKILLed, and that the container then shows exit code 137.
Explain the cgroup mechanism, reclaim before kill, cgroup-scoped versus global OOM killer, and the memory-plus-swap meaning of --memory-swap.
Diagnose from State.OOMKilled, memory.events and kernel logs, distinguish a killed child from a killed PID 1, and argue for disabling swap so failure is fast and observable.
Set policy: limits on every container so the host-wide OOM killer never runs, reservation-versus-limit conventions, and headroom sized from measured peaks per workload class.
## The mechanism `docker run --memory=512m` does one thing: it sets a memory limit on the control group (cgroup) that the container's processes live in — `memory.max` on cgroup v2, `memory.limit_in_bytes` on v1. A cgroup is the kernel's accounting and limiting unit for a set of processes; every page a process in that group touches is charged to it. A common misconception is that hitting the limit makes `malloc` fail. It usually does not. Linux overcommits: the allocator hands back an address range and only maps physical pages when they are first written. The reckoning therefore happens at page-fault time. When charging a new page would push the cgroup over its limit, the kernel first tries to reclaim — writing back and dropping page cache, evicting clean pages, swapping if swap is available to the cgroup. Only when reclaim cannot free enough does the **cgroup OOM killer** run. ## Cgroup OOM versus host OOM The cgroup OOM killer is scoped: it picks a victim from the processes in that cgroup, ranked by an OOM score roughly proportional to memory footprint, and sends SIGKILL. Nothing outside the container is touched — that is exactly the isolation you bought by setting a limit. The kernel logs a line such as `Memory cgroup out of memory: Killed process 1234 (java)`, visible through `dmesg` or the journal, and increments the `oom_kill` counter in the cgroup's `memory.events`. The **global** OOM killer is a different event: the whole host ran out, and the kernel picks a victim across all processes, which may be an unrelated container or a system daemon. Unlimited containers are the usual cause — the strongest argument for setting limits everywhere rather than nowhere. If the killed process is PID 1, the container's main process is gone, so the container stops. `docker inspect` then shows `State.ExitCode: 137` — 128 plus signal 9 — and `State.OOMKilled: true`. If the victim was a child (a forked worker, a shell), the container keeps running with a mysteriously missing subprocess, a far nastier failure to debug. SIGKILL cannot be caught, so there is no shutdown hook, no flush, no stack trace: the absence of application logs at the moment of death is itself the signal. ## Swap `--memory-swap` is the most misread flag in this area because it sets the **combined** memory-plus-swap ceiling, not the swap allowance. With `--memory=512m`: - `--memory-swap` omitted: total 1g, so 512m of swap is available. - `--memory-swap=512m` (equal to memory): swap effectively disabled; the container OOMs as soon as it exceeds 512m of RAM. - `--memory-swap=1g`: 512m RAM plus 512m swap. - `--memory-swap=-1`: unlimited swap, bounded only by host swap. Swap trades a kill for latency: the container survives but pages in and out, and throughput can collapse by orders of magnitude while everything still looks "up". For latency-sensitive services many teams deliberately disable swap so failure is fast and loud rather than slow and invisible. `--memory-swappiness` (0–100) tunes how eagerly the kernel swaps that container's anonymous pages. Swap accounting must be enabled in the kernel, or Docker warns and the swap settings are ignored. ## Related knobs - `--memory-reservation` is a **soft** limit: under host pressure the kernel tries to push the container back toward it, but it is not a kill boundary. Setting reservation below limit expresses "normally this much, occasionally more". - `--oom-kill-disable` stops the kernel killing anything in the cgroup; processes are frozen instead when the limit is hit. It is almost always a trap — a hung container that never recovers — and dangerous without a limit set. - `--oom-score-adj` biases which processes the *global* killer prefers. ## Practical takeaways Size the limit from measured peak, not average, plus headroom, because the limit is a hard cliff for an incompressible resource: unlike CPU, memory cannot simply be shared more slowly. Expect the OOM signature to be a silent restart loop with exit 137 and no application error, and confirm it from `docker inspect`, `memory.events` and kernel logs rather than guessing. And remember that everything the process touches counts toward the same cgroup limit — native buffers, thread stacks, page cache for files it writes — not just whatever the runtime reports as its own memory.
- The container is still running but a worker process vanished. Can that still be an OOM kill?Yes. The cgroup OOM killer picks the highest-scoring process in the cgroup, which is often a fat child rather than PID 1. The container stays up because its main process survived, so there is no exit code to inspect; you confirm it from the kernel log and the `oom_kill` counter in the cgroup's `memory.events`.
- Why does the application log show nothing at all before the container dies?SIGKILL cannot be caught, blocked or handled, so no signal handler, shutdown hook or log flush runs, and buffered lines die with the process. That silence combined with exit code 137 is a strong fingerprint of an OOM kill rather than an application crash.
- Is setting `--memory-swap` equal to `--memory` a good default?For latency-sensitive services usually yes: it disables swap for that container so an over-budget container fails fast and visibly instead of thrashing while appearing healthy. For batch jobs that spike occasionally, allowing some swap can be the cheaper trade. The decision is whether slow-and-alive or dead-and-restarted is the better failure mode.
A memory limit is a weight limit in a lift, not a queue: a CPU limit makes everyone move slower, but exceeding the memory limit means somebody is thrown out immediately.
saying these in an interview costs you the question
- Claiming allocation fails or throws at the limit, rather than the kernel killing at page-fault time
- Thinking `--memory-swap` sets the swap size instead of the memory+swap total
- Assuming an OOM kill inside one container takes down other containers or the host
- Expecting a graceful shutdown, stack trace or log flush after SIGKILL
- Reaching for `--oom-kill-disable` to 'fix' OOM kills