How do you use `docker events --filter` to catch containers that die overnight?
answer
- The stream starts when you start it
- Replay depends on a bounded memory buffer
- Same key ORs, different keys AND
- One action carries the exit code
- Automatic cleanup erases the post-mortem
basics
~20 sRun docker events filtered to the actions you care about — --filter event=die --filter event=oom --filter event=health_status — and capture it to a file, because the stream is live and the daemon's replay buffer is small. Each die event carries the container's exitCode as an attribute.
solid answer
~50 s`docker events` is the daemon's live activity stream: every object it touches emits an event with a type, an action and an actor carrying attributes. For containers that vanish overnight the useful actions are `oom`, `die` (whose attributes include `exitCode` and the image and name), `kill`, `stop`, `destroy` and `health_status`. Filter with repeated `--filter` flags — `--filter type=container --filter event=die` — and note that repeating the *same* key ORs the values while different keys AND together. `--since`/`--until` replay from the daemon's in-memory buffer, which is bounded and does not survive a daemon restart, so the reliable pattern is to run the command as a long-lived capture writing `--format '{{json .}}'` to a file before the next occurrence. When something does die, the events give you the sequence and the timestamp; `docker inspect` on the container (if it survives) gives you the final state. With `--rm` containers there is no post-mortem at all, and the event stream is your only record.
code
bash · 15 lines# Live stream, narrowed to the deaths and health transitions of containers
docker events --filter type=container \
--filter event=oom --filter event=die --filter event=health_status
# Same keys repeated are ORed; different keys are ANDed
docker events --filter type=container --filter event=die --filter event=kill
# Replay from the daemon's in-memory buffer (bounded, lost on daemon restart)
docker events --since 30m --until 5m --filter event=die
# Durable capture: one JSON object per line, appended to a rotated file
docker events --filter type=container --format '{{json .}}' >> /var/log/docker-events.jsonl
# Narrow to one workload by label
docker events --filter 'label=app=thumbnailer' --filter event=diego deeper
Know that the daemon emits a stream of events and that docker events shows it live; be able to name a few container actions such as create, start, die and destroy.
Explain the record's shape — type, action, actor attributes — the filter keys, and the AND/OR rule for repeated filters, plus why --since replay is limited.
Show the incident method: start a capture before the next occurrence, read the ordered sequence, and pair it with the container's own record; recognise that automatic cleanup destroys the evidence.
Own the operational stance — that a live-only stream is not history, and that a cheap per-host capture with rotation is what turns overnight mysteries into timelines across a fleet.
## The stream `docker events` subscribes to the daemon's event bus and prints everything the engine does as it happens. Each record has a **type** (`container`, `image`, `network`, `volume`, `daemon`, `plugin`), an **action** (`create`, `start`, `die`, `kill`, `stop`, `destroy`, `oom`, `health_status`, `pull`, `connect`, …), an **actor** (the object's ID plus attributes such as `name`, `image` and the object's labels) and a timestamp. It is push, not poll: nothing is aggregated and nothing is retried, you simply see what the daemon did. The critical property is that **it is live**. Start it now and you see events from now on. `--since` and `--until` replay from a buffer the daemon keeps in memory, which is bounded and is gone when dockerd restarts — so it is a convenience for "what happened in the last few minutes", never an audit log. ## Filtering Filters are repeated `--filter key=value` flags. The keys worth memorising are `type`, `event` (the action), `container`, `image`, `label` and `daemon`. The combination rule catches people out: repeating the *same* key ORs its values, while *different* keys AND together. So `--filter event=die --filter event=oom` shows both actions, but `--filter type=container --filter event=die` narrows to container deaths only. For a container that keeps disappearing, the actions to watch are: - **`oom`** — the kernel OOM-killed a process in the container. This is emitted separately from the death itself, so seeing `oom` immediately before `die` is the strongest signal you will get from the engine that memory, not the application, ended it. - **`die`** — the container's main process exited. The attributes include `exitCode`, `image` and `name`, which is why `docker events` is often faster than reconstructing the story from anywhere else. - **`kill`** — a signal was delivered, with the signal in the attributes. - **`stop`**, **`destroy`** — a deliberate stop, and removal of the container record (which is what `--rm` triggers right after `die`). - **`health_status`** — emitted on every transition of a HEALTHCHECK's verdict, so a container drifting unhealthy before it dies shows up here first. `--format` accepts a Go template over the record, and `--format '{{json .}}'` is the form to use for capture because it emits one self-describing JSON object per line that you can grep or post-process later. ## A worked incident A 47-node CI fleet runs a Node.js worker that consumes a queue and produces image thumbnails. Overnight, roughly 23 jobs fail with no application error and no container to inspect — the workers are started with `--rm`, so by morning the containers are gone and `docker ps -a` shows nothing at all. Because nothing was capturing events, the first move is to *start* capturing on a few nodes and wait for the next occurrence: Run a long-lived `docker events --filter type=container --filter event=oom --filter event=die --filter event=health_status --format '{{json .}}'` appending to a file. The next night the file shows, for one worker, a `health_status` transition to unhealthy, then an `oom` event, then `die` with an exit code attribute, then `destroy` a fraction of a second later — the `--rm` cleanup that erased the evidence. The sequence localises the problem to memory pressure inside the thumbnail worker rather than to the queue, the image or the scheduler, and it does so without reproducing anything. The follow-up is to stop the evidence disappearing: drop `--rm` for the workload while it is under investigation so the container record survives for `docker inspect`, which then gives you the final `State` — including whether the engine recorded the container as OOM-killed and when it started and finished. ## Practical notes **Capture before you need it.** Because there is no durable history, the single highest-value operational habit around this command is to run one capture per host as a supervised service, writing JSON lines to a file with normal log rotation. It costs almost nothing and it converts "we have no idea what happened at 3am" into a timestamped sequence. **Events are per daemon.** There is one stream per engine; on a fleet you either capture on each host or point a client at each engine. There is no cross-host aggregation in the command itself. **Health events are notifications, not actions.** A plain engine does not restart a container because its health check went unhealthy — it records the transition and emits the event. Something above the engine has to act on it. Knowing that distinction is often the real point of the question. **Pair it with the record, not instead of it.** Events tell you the *sequence* and the timestamps; the container's own record tells you the *final state* and the configuration that produced it. Investigations that use only one of the two either lack the timeline or lack the detail.
- Why can't you use `docker events --since 8h` the morning after an overnight failure?Because the replay comes from a bounded in-memory buffer in the daemon that is discarded when dockerd restarts, so a window that long usually returns nothing useful. The command is a live subscription first and a short replay second. The fix is to run a capture continuously — `docker events --format '{{json .}}'` appended to a rotated file — so the history exists before you need it.
- A container's HEALTHCHECK goes unhealthy and emits a health_status event. Does the engine restart it?No. A plain Docker Engine records the transition and emits the event; it does not act on it. A restart policy reacts to the container's *process exiting*, not to a health verdict, so an unhealthy container that keeps running stays running. Acting on unhealthy is the job of whatever supervises the container above the engine.
- How do the event stream and `docker inspect` divide the work in an investigation?Events give you the ordered timeline with timestamps — health transition, oom, die with its exit code, destroy — including for containers that no longer exist. `docker inspect` gives you the surviving container's final `State` and the full configuration that produced it. Use events to establish what happened and when, then inspect to establish what the container actually was.
- What do you lose by running short-lived containers with `--rm` while debugging?The container record is destroyed moments after the process exits, so there is nothing left to inspect — no final State, no exit code on disk, no writable layer to examine. The event stream is then your only trace. While a failure is under investigation, drop `--rm` so the stopped container survives, and clean up explicitly afterwards.
saying these in an interview costs you the question
- Believes the event stream is a durable audit log
- Expects --since to replay hours of history
- Thinks an unhealthy health check restarts the container
- Cannot name any event action besides start and stop
- Assumes filters always AND together
- Keeps --rm on a workload under investigation