A container on a Docker host keeps disappearing and nobody sees it happen. Describe how you would use 'docker inspect' and 'docker events' to establish what is actually going on.
answer
- inspect = state snapshot, events = timeline
- --format '{{.State.ExitCode}}' / .State.OOMKilled
- events --since/--until replays a window
- filter container= / event=die
- events are in-memory, not persisted
basics
~20 sdocker inspect dumps a container's full JSON state — exit code, OOMKilled flag, restart count, mounts, network, config — for the current or last run. docker events streams the daemon's real-time lifecycle events (create, start, die, kill, destroy, health_status) so you can watch it happen and correlate timings.
solid answer
~50 s`docker inspect <c>` is the post-mortem: `--format` pulls out `.State.ExitCode`, `.State.OOMKilled`, `.State.Error`, `.State.FinishedAt`, `.RestartCount`, plus the effective config — mounts, env, log driver, resource limits, networks. It answers 'how did it end and with what settings', and it works on a stopped container as long as it has not been removed. `docker events` is the live view: the daemon streams every lifecycle transition with timestamps — `create`, `start`, `die` (with the exit code attribute), `kill`, `oom`, `health_status`, `destroy` — and events for images, volumes and networks. `--since`/`--until` replay a past window and `--filter` narrows to one container or event type, which is how you catch something that vanishes when nobody is watching. The two together usually settle it: events show the sequence and who acted, inspect shows the resulting state. Add `docker logs` for what the application itself said just before.
code
bash · 6 linesdocker inspect -f '{{.State.ExitCode}} oom={{.State.OOMKilled}} restarts={{.RestartCount}}' web
docker inspect -f '{{json .State.Health}}' web | jq
docker events --since 2h --filter container=web \
--format '{{.Time}} {{.Action}} {{.Actor.Attributes.exitCode}}'
docker events --filter event=oom --filter event=diego deeper
Know that inspect prints a container's full configuration and last state, and that events streams lifecycle transitions live.
Extract specific fields with --format, and use --since/--until plus filters to replay a past window.
Combine timeline, state and logs into one causal story, and know the limits — no persistence for events, no application-internal visibility.
Insist on continuous event and log capture off-host so post-incident reconstruction does not depend on a container that has already been recreated.
## docker inspect: the state snapshot `docker inspect` returns the daemon's complete JSON view of an object. For a container that includes: - **`.State`** — Status, Running, Pid, `ExitCode`, `Error`, `OOMKilled`, StartedAt, FinishedAt, and `Health` when a health check is defined (status plus the last probe outputs). - **`.RestartCount`** and `.HostConfig.RestartPolicy` — whether something is restarting it. - **`.HostConfig`** — resource limits, log driver and options, capabilities, privileged flag. - **`.Config`** — image, entrypoint, cmd, env, user, labels. - **`.Mounts`** and **`.NetworkSettings`** — what is actually mounted and which networks and IPs are attached, which is where 'it works on my machine' arguments usually die. Use `--format` with a Go template to extract precise fields, or `--format '{{json .State}}'` piped to `jq`. The key operational point: inspect reports the *effective* configuration the daemon applied, not what someone believes the compose file says, so it is the fastest way to catch drift between intent and reality. Inspect works on stopped containers too, so an exit code and OOMKilled flag survive until the container is removed. If a restart policy or an orchestrator removes and recreates the container, that evidence is destroyed — which is exactly the case where events matter. ## docker events: the timeline `docker events` streams daemon-level events as they happen. Container events include `create`, `start`, `restart`, `die` (carrying `exitCode`), `kill` (carrying the signal), `stop`, `oom`, `destroy`, `health_status`, plus `exec_create`/`exec_start` and image, volume, network and plugin events. Useful invocations: - `docker events --filter container=web --filter event=die` - `docker events --since 1h --until 10m` to replay a window after the fact - `--format '{{.Time}} {{.Type}} {{.Action}} {{.Actor.Attributes.name}}'` for a compact feed The replay ability is what makes this the right tool for intermittent disappearance: you get a timestamped ordering showing, for example, `die` with exit code 137 immediately preceded by an `oom` event, or `kill` with signal 15 followed by `destroy`, which points at an operator or a deployment tool rather than the app. ## Putting it together A workable sequence for 'the container keeps vanishing': 1. `docker events --since 2h --filter container=<name>` — build the timeline. Is it dying on its own, being killed, or being destroyed and recreated? 2. `docker inspect` the current or last instance — exit code, OOMKilled, Error, RestartCount, and the configured restart policy and memory limit. 3. `docker logs --tail 200 --timestamps` — what the application printed just before the last transition. 4. Correlate: an `oom` event plus `OOMKilled: true` plus exit 137 is a memory limit problem; a clean exit 0 with a restart policy of `always` is a service that finished its work; repeated `die` shortly after `start` with a nonzero code is a startup failure to read in the logs. ## Practical cautions - Events are **not persisted**. The daemon keeps a bounded in-memory buffer, so `--since` reaches only so far back; on a host you care about, stream events to a file or a collector continuously. - `docker inspect` output is large and its shape varies across engine versions — always target fields with `--format` rather than parsing positions. - Neither tool sees inside the application. They tell you what the daemon did and what state resulted; the application's own logs and metrics tell you why.
- docker events shows an 'oom' event and inspect reports exit code 137 with OOMKilled true. What does that combination mean?The kernel's out-of-memory killer terminated the process because the container exceeded its memory limit, and 137 is 128 plus signal 9, the SIGKILL that resulted. It points at either a limit set too low for the workload or genuine excess allocation in the application, and it is settled by comparing the limit with observed usage rather than by raising the limit reflexively.
- Why might docker events show nothing for an incident that happened yesterday?The daemon keeps events in a bounded in-memory buffer rather than persisting them, so --since can only reach back as far as that buffer holds and everything before it is gone. On hosts that matter, stream docker events to a file or a log collector continuously so past incidents remain reconstructable.
saying these in an interview costs you the question
- Believing docker events are stored permanently and queryable forever
- Reading container state from the compose file instead of docker inspect
- Assuming inspect is unusable once the container has stopped
- Parsing full inspect JSON by position instead of using --format
- Expecting these tools to explain application-internal causes