A Docker container has restarted 47 times overnight — what evidence about the earlier runs survives?
answer
- The container object is reused
- One stream, many runs
- State is a snapshot, not a history
- One counter is the only aggregate
- Rotation decides how far back you see
basics
~20 sA restart policy reuses the same container, so docker logs still holds every earlier run's output, separated only by timestamps. docker inspect describes just the latest run apart from RestartCount, and log rotation may have deleted the first failure.
solid answer
~40 sThe key fact is that a Docker restart policy puts **the same container object** back, keeping its ID and its log file. So `docker logs <container>` contains the output of all 47 runs concatenated, with no separator between them — use `--timestamps` and `--since` to slice out one attempt, since each run's startup banner marks a boundary. `docker inspect`, by contrast, is not a history: `.State.ExitCode`, `.State.OOMKilled`, `.State.StartedAt` and `.State.FinishedAt` all describe the latest run only, and `.RestartCount` is the sole aggregate. What can be missing is the beginning: if the json-file driver is configured with `max-size` and `max-file`, output older than that window has been rotated away, so the *first* failure — usually the informative one — may be gone. Capture logs to a file early in an incident.
code
bash · 3 linesdocker logs --timestamps pdfsign-api > /tmp/pdfsign-crashloop.log
docker inspect --format '{{.RestartCount}} {{.State.ExitCode}} {{.State.OOMKilled}}' pdfsign-api
docker logs --timestamps --since 2026-08-19T02:00:00Z --until 2026-08-19T02:05:00Z pdfsign-apigo deeper
Remember that a Docker restart policy restarts the same container, so docker logs on it still contains the output of earlier attempts. Adding --timestamps is what makes that pile of output readable.
Explain why logs accumulate but inspect does not: one log stream per container object versus a State snapshot overwritten at every start. Know that RestartCount is the only aggregate the engine keeps.
Show the incident reflex — capture the full log to a file before anything rotates or prunes it, segment by timestamp, and read the earliest surviving run rather than the newest, because later failures are usually consequences.
Own the retention question: rotation sizing that keeps a loop from filling the disk without erasing the first failure, log shipping so evidence outlives the host, and a norm that nothing is removed before it is captured.
When a container under a restart policy has been failing all night, the useful question is not "what happened?" but "what evidence still exists, and what has already been destroyed?". Docker's answer is unusual enough to be worth knowing precisely, because it differs from the intuition people bring from orchestrators. ## The same container comes back A Docker restart policy does not create a replacement container. The daemon restarts **the same container object**: same ID, same name, same writable layer, same log file. Three consequences follow, and they define what triage is possible. **Logs accumulate.** Because the json-file driver writes to a path owned by that container, output from every run appends to one stream. `docker logs pdfsign-api` therefore returns all 47 runs, in order, with nothing marking where one ended and the next began. This is a gift and a trap: the history is there, but it reads as one confusing narrative unless you segment it. `--timestamps` gives every line an RFC3339 prefix; the application's own startup banner is the natural run boundary; `--since` and `--until` accept timestamps or relative durations, so `docker logs --timestamps --since 2026-08-19T02:00:00Z --until 2026-08-19T02:05:00Z pdfsign-api` isolates one window. `--tail` counts from the end, which by definition lands you in the most recent attempt. **The writable layer accumulates too.** Files the application wrote during earlier runs are still there unless it truncates them at startup, so `docker cp` can retrieve a crash file written hours ago. This also means disk use grows across a long crash loop. **`docker inspect` does not keep history.** `State` is a snapshot of the latest run: `ExitCode`, `OOMKilled`, `Error`, `StartedAt` and `FinishedAt` are overwritten every time. There is no list of previous terminations. The one aggregate is the top-level `RestartCount`, and even that is a count, not a record. So if run 12 was killed for memory and runs 13 to 47 died of something else, `inspect` will only ever tell you about run 47 — the logs, not the state, hold the story. ## What has already been destroyed **Rotation.** The json-file driver takes `max-size` and `max-file` options, either per container (`--log-opt max-size=10m --log-opt max-file=3`) or as daemon-wide defaults in `daemon.json`. Rotation is what stops a crash loop filling the disk, and it is also what deletes the earliest attempts. `docker logs` reads the rotated files the driver still keeps, but anything beyond `max-size` times `max-file` is gone. In a loop that restarts every few seconds and logs a traceback each time, a 10 MB window can be consumed in minutes — which is exactly how teams lose the *first* failure, the one that was not yet contaminated by downstream effects. **Removal.** `docker rm` on the container, or a `docker system prune` that sweeps stopped containers, destroys logs and inspect record together. During an incident, copy first: `docker logs --timestamps pdfsign-api > pdfsign-crashloop.log`. **A non-local driver.** If the container's logging driver ships output off the host, the local read may be limited to a small cache and the real history lives in the destination system. Check with `docker inspect --format '{{.HostConfig.LogConfig.Type}}' <container>`. ## A worked example A PDF-signing service — a Django application under gunicorn — is looping. `docker inspect --format '{{.RestartCount}} {{.State.ExitCode}} {{.State.OOMKilled}} {{.State.StartedAt}}' pdfsign-api` shows `47 1 false 2026-08-19T06:12:09Z`: the latest run failed with a plain non-zero status, no out-of-memory involvement. That is only run 47. Dumping `docker logs --timestamps pdfsign-api > /tmp/loop.log` and counting startup banners shows 47 boot sequences, and the *first* one — still inside the rotation window by luck — contains a stack trace about an unreachable signing-key service that no later run repeats, because after the first failure the app fails earlier, on a stale lock file it left behind. Without the accumulated log, the investigation would have chased the lock file and never found the trigger. The lesson generalises: in a loop, the earliest surviving evidence is usually worth more than the latest. ## Live observation `docker events --filter container=pdfsign-api` streams the daemon's own view — `start`, `die` with its exit status, `oom`, `restart` — as it happens, and `--since` replays recent history the daemon still holds. That gives you an authoritative timeline of transitions to line up against the log timestamps, which is particularly useful when the application's clock or log format makes its own ordering unclear. ## The habit In any crash loop: capture the whole log to a file before anything can rotate or prune it, segment it by timestamps into runs, read the *earliest* one you still have, and treat `inspect` as a description of the last attempt only. Then check whether rotation settings are quietly deciding how much of your evidence exists.
- What limits how far back `docker logs` can go on a container that has been looping for hours?The json-file driver's rotation settings. `max-size` bounds each file and `max-file` bounds how many are kept, so the readable window is roughly their product; anything older has been deleted. These can be set per container with `--log-opt` or as daemon defaults in `daemon.json`. With a driver that ships logs off the host, the local window may be a small cache instead, and the real history lives in the destination.
- How do you split one accumulated log stream back into individual runs?Use `docker logs --timestamps` so every line carries an RFC3339 prefix, then cut on the application's own startup banner, which marks each new run. `--since` and `--until` narrow the dump to one window. Cross-check the boundaries against `docker events --filter container=<name>`, which gives the daemon's `start` and `die` transitions independently of anything the application prints.
- Why can `docker inspect` be actively misleading during a long crash loop?Because `State` is overwritten on every restart, so it describes only the most recent run. A loop that began with an out-of-memory kill and later degraded into a different failure will show only the latter, and the top-level `RestartCount` is the sole aggregate. The accumulated log — read from its earliest surviving entry — is where the sequence lives.
saying these in an interview costs you the question
- Assumes each restart creates a new container with fresh logs
- Reads only the tail and misses the first failure
- Treats inspect State as a history of past exits
- Runs docker rm or a prune before capturing evidence
- Ignores log rotation limits when history looks suspiciously short
- Cannot say where RestartCount comes from