How do you build a Docker Engine API /events client that misses no container deaths?
answer
- The stream is a signal, not a log
- It will end; plan for that
- There is a parameter that backfills
- Pair it with a full list call
- Subscribe before listing, then deduplicate
basics
~20 sTreat GET /events as a long-lived stream that will break: reconnect with backoff, pass since set to the last event's time so the daemon replays the gap, and reconcile against GET /containers/json on every reconnect because that replay buffer is bounded.
solid answer
~50 s`GET /events` is a chunked HTTP response the daemon holds open, emitting one JSON object per event with `Type`, `Action`, `Actor.ID` and `Actor.Attributes`; filters are passed server-side as a URL-encoded JSON map so you are not parsing the whole firehose. Building a correct client means accepting three facts. The stream **will** end — a daemon restart, a reload, a proxy timeout or a network blip closes it — so EOF is normal and the loop reconnects with backoff rather than exiting. The stream is **edge-triggered**, so on reconnect you pass `since` with the timestamp of the last event you processed and the daemon replays the gap from a bounded in-memory buffer; a long outage still leaves a hole. Therefore you also **reconcile**: open the stream first, then call `GET /containers/json?all=true` and diff it against your own state. Keep handlers idempotent, since replay re-delivers events.
code
bash · 3 linescurl -sN -G --unix-socket /var/run/docker.sock \
--data-urlencode 'filters={"type":["container"],"event":["die"]}' \
http://localhost/v1.43/eventsgo deeper
Know that the Docker Engine API can push events rather than being polled, and that GET /events streams one JSON object per container start, die or destroy. Writing the resilient client is not expected yet.
Be able to describe the stream concretely: a chunked response, one JSON object per event with type, action and actor, filters encoded as a JSON query parameter, and since/until to bound or replay a window.
Demonstrate that you have run one of these: reconnect on EOF with backoff, backfill with since while knowing the buffer is bounded, reconcile against a list call in the right order, and keep handlers idempotent.
Own the pattern rather than the endpoint — edge-triggered notification plus a level-triggered reconciliation loop — and set the expectation that any agent driving the engine is designed to converge rather than to trust a stream.
### What the endpoint is `GET /events` is the Docker Engine API's notification stream. The daemon answers with a chunked response it never ends on its own and writes one JSON object per event as things happen: a container is created, starts, dies, is destroyed; an image is pulled or removed; a volume is mounted; a network is connected. Each object carries a `Type` (`container`, `image`, `network`, `volume`, `daemon`), an `Action` (`start`, `die`, `destroy`, …), an `Actor` with the object's `ID` and an `Attributes` map — for a container death that map includes the image, the container name and the exit code — and both a second-resolution `time` and a nanosecond `timeNano`. Filtering is done on the daemon side by URL-encoding a JSON map into the `filters` query parameter, so a watcher that only cares about container deaths never has to receive image pulls at all. `since` and `until` bound the window: `until` turns the call into a finite historical query, and `since` alone means "replay from that instant, then keep streaming". ### The failure this question is really about Here is the incident. A supervisor process watches container deaths for a fleet of PDF-signing workers and re-queues the job of any worker that dies. It opens `/events`, filters to `die`, and works perfectly for weeks. Then the host's nightly maintenance restarts `dockerd`. The stream ends; the supervisor's read loop sees EOF, reconnects successfully, and carries on — but the 47 workers that died during and just after the restart were never reported, so 47 jobs sit in the queue marked *in flight* until a human notices. Nothing in that story is a Docker bug. An event stream is an **edge-triggered** signal: it tells you about transitions while you are listening, and it owes you nothing about transitions that happened while you were not. Every correct client is built around that. ### Three rules for a correct watcher **1. Expect the stream to end.** A daemon restart, a config reload, an idle timeout in anything sitting between you and the daemon, or a plain network fault will close it. EOF is a normal condition, not a crash: log it, back off, reconnect. A watcher that exits on EOF and relies on a supervisor to restart it can work, but only if the restart carries the state needed for the next two rules. **2. Backfill with `since`.** On reconnect, pass `since` set to the timestamp of the last event you actually processed — the nanosecond field is there so you can be precise — and the daemon replays what it holds from that moment before resuming live streaming. That closes short gaps completely. It does not close long ones: the daemon's event history is a bounded in-memory buffer, not a durable log, and it does not survive the daemon restarting. Never treat `since` as a guarantee. **3. Reconcile on every connect.** Because the backfill is best-effort, pair the edge-triggered stream with a level-triggered check: after the stream is open, call `GET /containers/json?all=true` and diff the result against the state you believe in. Anything you think is running that is not there, or is there with an exit code you never processed, gets handled now. The ordering matters and is the part people get wrong — **open the stream first, then list**. Do it the other way round and a container that dies in the window between the list returning and the subscription taking effect is lost forever, which is the same bug in a smaller window. ### Two more things a production client gets right **Idempotency.** Replay from `since` can re-deliver events you have already handled, and the boundary event is easy to see twice. Key your handler on something stable — container ID plus action plus `timeNano` — and make re-handling a no-op. A supervisor that re-queues a job every time it sees the same `die` is worse than one that misses it. **Never work inside the read loop.** Decode the event, push it onto a queue, and return to reading immediately. Doing a database write or an HTTP call inline means you are not reading the stream while it happens, which is exactly when you want to be reading it. Deserialise defensively, too: new event types and new attributes appear over time, and an unknown `Action` should be ignored rather than fatal. ### Testing it The test that matters is not "does it see a `die`" — it is "does it recover". Restart the daemon underneath the watcher; kill the connection mid-stream; suspend the process for longer than the daemon's history and confirm the reconciliation pass, not the backfill, is what repairs the state. A watcher that has never been tested against a daemon restart has not been tested. ### What an interviewer is testing Whether you have run one of these in production. The junior answer is "open /events and read it". The senior answer names the edge-versus-level distinction, uses `since` while knowing its limits, reconciles against a list call in the right order, and makes the handler idempotent.
- Your watcher reprocesses events after a reconnect. Is that the daemon misbehaving?No. Replaying from `since` re-delivers everything at and after that instant, so the boundary event is commonly seen twice, and any overlap between the backfill and your reconciliation pass produces duplicates by design. The client's job is idempotency: key on container ID plus action plus the nanosecond timestamp and make a repeat handling a no-op.
- Why open the event stream before listing containers rather than after?To close the race. With the stream already open, anything that happens while the list call is in flight still reaches you, and you drop it as a duplicate. List first and subscribe second, and a container that dies in that window appears in neither the list nor the stream — the exact bug the reconciliation was meant to prevent, just with a shorter window.
- What ends an event stream besides a network fault?A daemon restart or a reload of its API server, an idle or read timeout in anything sitting between client and daemon, and the client's own read timeout if it sets one on a stream that is legitimately silent for hours. Treat all of them as ordinary: reconnect with backoff and jitter rather than exiting or hammering.
The event stream is a doorbell, not a guest list. It tells you when someone arrives while you are home; when you get back from an errand you still have to look around the house.
saying these in an interview costs you the question
- Assumes the stream stays open indefinitely
- Reconnects without a since parameter
- Never reconciles against the actual container list
- Lists containers first, then subscribes
- Does slow work inside the read loop
- Treats duplicate events as impossible