skip to content

How would you choose a fleet-wide container logging strategy: a shipping log driver in dockerd, or files plus a host collector?

level: principalimportance: should knowfreq 33%

answer

  1. Decide on coupling, not on convenience
  2. Ask what a log-backend outage costs you
  3. One path keeps the write purely local
  4. State the loss semantics in one sentence
  5. Configuration is frozen when a container is created

basics

~20 s

Choose on failure coupling. A shipping driver inside dockerd puts a remote system on every container's write path; bounded local files read by a host collector keep that write local and logs readable on the host, at the cost of an agent.

solid answer

~50 s

There are three shapes: a **shipping driver** in the daemon (fluentd, gelf, awslogs), **bounded local files** (`local` or `json-file` with rotation) read by a collector on the host, or the application shipping its own logs. The deciding question is *what happens when the log backend is unavailable*. With a shipping driver the answer reaches the container's `stdout` write — blocking mode stalls the app, non-blocking drops messages — so a log outage becomes an availability event. With files plus a collector, the container only ever writes to local disk; the collector falls behind and catches up, and rotation bounds the exposure. Files also keep `docker logs` useful for on-call and let you set one host disk budget. I would make bounded `local` the fleet default in `daemon.json`, allow narrow per-container overrides, and put durability in the collector tier rather than in the engine.

go deeper

for a junior

You are unlikely to be asked to own this, but know the two shapes exist: logs written to files on the host, or sent straight out by the daemon to a remote destination.

for a middle

Be able to compare the options concretely — where the bytes go, what has to be installed, and what accumulates on the host — and to explain why bounding local files is not optional.

for a senior

Argue the failure coupling: a shipping driver puts the log backend on every container's write path, while files plus a collector turn a backend outage into delayed logs. Bring the disk arithmetic and the on-call read path.

for a principal

Own the policy end to end: one default enforced in daemon.json, a written statement of loss semantics, buffering and redaction in a collector tier you can change centrally, and a rollout plan that accounts for configuration being frozen at container creation.

## Frame the decision as coupling, not plumbing Every option moves the same bytes to the same place. What differs is **what the container depends on to keep running**, and that is the axis a lead should decide on. **Option A — a shipping driver in the daemon.** Each container's messages go straight from `dockerd` to a remote endpoint. Nothing accumulates on the host, records are structured and tagged with container metadata at the source, and there is no agent to deploy. The price is that the log backend is now on the write path of every container: in the default blocking mode a stalled endpoint back-pressures the application's writes to `stdout`; in non-blocking mode you instead lose messages once the buffer fills. You have made a support system a hard dependency of production traffic. **Option B — bounded local files plus a collector.** Containers log through `local` or `json-file` with `max-size` and `max-file` set; a collector process on the host tails those files and forwards them. The application only ever writes to local disk. If the backend disappears, the collector queues, retries and catches up; the only real exposure is rotation discarding lines the collector never got to, which is a function of the bound you chose. You pay for an agent per host and for disk. **Option C — the application ships its own logs.** Direct, but it puts retries, batching and credentials into every service in every language, and it usually loses the console output that everything else — `docker logs`, crash output, panics from the runtime itself — still goes through. Worth it only for a genuinely special stream. ## The criteria I would actually weigh 1. **Blast radius of a log-backend outage.** With Option A it is the fleet; with Option B it is delayed logs. This alone decides most cases. 2. **On-call ergonomics.** Can an engineer read a container's recent output on the host without the log platform? File-based drivers answer natively. Shipping drivers answer only through the daemon's local read cache, and only if nobody disabled it. 3. **Loss semantics you can state out loud.** "We may drop logs during a backend outage" or "we may stall" or "we keep the last N megabytes on disk" — the fleet should have one written answer, not a per-team accident. 4. **Host disk budget.** With files this is explicit: per-container bound × containers per host, plus headroom on the filesystem that also holds images and volumes. Do the arithmetic once and enforce it as the default, because a bound that must be remembered on each `docker run` will be forgotten. 5. **Change safety.** Log configuration is frozen at container creation, so a fleet-wide change to `daemon.json` only reaches workloads as they are recreated. Any migration is a rolling one, and for a period you are running both shapes. 6. **Portability.** File-based logging with a node collector is the shape orchestrators use as well, so the collector tier and the pipeline survive a platform migration; a driver-specific configuration on every container does not. 7. **Metadata and cost.** Shipping drivers can attach container labels and environment values at the source, which is convenient but also the point where volume and cardinality get expensive. That budget is easier to enforce in a collector you control than in a driver option repeated across teams. ## The position I would defend Make **bounded local files the default**, set in `daemon.json` so a container created by anyone gets it, with a per-container bound derived from the host disk budget. Run one collector per host to forward and enrich, and treat *that* tier as the place where buffering, retry, sampling and redaction live — it can be restarted, upgraded and rate-limited without touching a single workload. Keep `docker logs` working so that incident response does not require the log platform to be healthy. Allow a shipping driver as a **narrow, per-container exception**, with `mode=non-blocking` and a buffer sized from that service's measured log rate, for the rare stream that genuinely must leave the host immediately. Write down the loss semantics for that exception. The answer that should be challenged in a review is "we ship straight from the driver because it is simpler". It is simpler, right up to the first log-backend incident, at which point the simplification is paid for in application availability — and the team discovers it during the outage rather than during the design.

  • How would you roll a fleet from unbounded json-file to a bounded default without a maintenance window?
    Set the new `log-driver` and `log-opts` in `daemon.json` so every newly created container picks them up, then let normal deploys convert the fleet, since log configuration is frozen at creation and neither a daemon restart nor `docker restart` re-applies it. For long-lived containers that are rarely redeployed, schedule recreation explicitly. Track progress by inspecting each container's actual LogConfig rather than assuming the daemon default reflects reality.
  • What would make you accept a shipping driver as the fleet default after all?
    Hosts with too little disk to hold a useful local window, a hard requirement that logs leave the machine before it can be destroyed, or a log path engineered to be more available than the workloads themselves. Even then I would run non-blocking with a measured buffer, keep the daemon's local read cache enabled so on-call can still read a container, and state the loss semantics explicitly.
  • Where should redaction of sensitive fields live in this design?
    In the collector tier, not in the engine. The log driver has no notion of fields it should not forward, its options are set per container by whoever wrote the run command, and changing them means recreating containers. A collector is one component you control, can update centrally, and can test — and it is also the natural place for sampling and volume caps, which are the same class of policy.

saying these in an interview costs you the question

  • Picks a driver on convenience without considering failure coupling
  • Treats the log backend as always available
  • Leaves the per-container bound to each team's run command
  • Assumes a daemon.json change reconfigures running containers
  • Has no stated answer for what happens to logs during an outage
  • Puts retries and redaction inside every application

context