skip to content

What does Docker's log-opt mode=non-blocking change, and what does max-buffer-size control?

level: seniorimportance: nice to knowfreq 26%

answer

  1. The default couples the app to logging
  2. A stalled driver reaches the application's write
  3. One option inserts a buffer in between
  4. When that buffer fills, something is discarded
  5. Its size defaults to one megabyte

basics

~20 s

Docker log drivers deliver in blocking mode by default, so a stalled driver back-pressures the container's writes to stdout and can stall the application. mode=non-blocking inserts a memory ring buffer, sized by max-buffer-size (1 MB default), that drops messages instead.

solid answer

~50 s

By default a Docker log driver delivers **blocking**: the daemon reads a message from the container's `stdout` and does not read the next one until the driver has accepted it. If the driver is slow — a stalled log endpoint, a full disk — the pipe fills and the *application's own write to stdout blocks*. That turns a logging outage into an application outage, which is the classic way a container that depends on nothing but a log collector goes unresponsive. Setting `--log-opt mode=non-blocking` inserts an in-memory ring buffer between the container and the driver: writes always complete, and if the buffer fills, log messages are dropped rather than the application being stalled. `--log-opt max-buffer-size` sizes that buffer, defaulting to `1m`. It is an explicit trade of log completeness for application availability, and it applies to any driver, not only the shipping ones.

code

bash · 6 lines
bash
docker run -d --name tiles \
  --log-driver=fluentd \
  --log-opt fluentd-address=logs.internal:24224 \
  --log-opt mode=non-blocking \
  --log-opt max-buffer-size=4m \
  tileserv:2.4

go deeper

for a junior

It is enough to know that writing to stdout inside a container is not free, and that Docker has a delivery mode option that decides whether a slow log destination waits or drops messages.

for a middle

Explain the mechanism: the daemon stops draining the container's stdout pipe until the driver accepts, so back-pressure reaches the application's write, and non-blocking mode inserts a bounded ring buffer instead.

for a senior

Show you would recognise the symptom in an incident — stalled writes to stdout that track the log backend's outage — and that you can size max-buffer-size from a measured log rate rather than guessing.

for a principal

Own the fleet-wide default: which classes of workload may lose logs to stay up, which must block, and the fact that leaving the default in place makes the log backend a hard dependency of every service.

## The failure mode A geospatial tile service — Django under gunicorn — shipped its logs with a remote driver. The log endpoint stalled for about forty-seven seconds during a maintenance window. In that window the tile service stopped answering: workers that would normally turn a request round in milliseconds were parked inside a `write()` to `stdout`, the health endpoint timed out, and instances were recycled with the usual six-second cold start each. Nothing was wrong with the application, the database, or the network path to the users. The log driver had applied back-pressure and the application had inherited it. This is the default behaviour, and it is a deliberate one: **blocking** delivery means no message is silently lost. ## How blocking delivery propagates The container's `stdout` is a pipe read by the daemon. The daemon takes a message off that pipe and hands it to the logging driver; only when the driver accepts it does the daemon come back for the next. If the driver is slow to accept, the daemon stops draining, the pipe's kernel buffer fills, and the next `write(2)` from the process inside the container blocks. From the application's point of view, printing a line has become a blocking network operation against a system nobody thought of as a dependency. The severity depends on the concurrency model. A worker-per-request server stalls one worker per blocked write and can wedge its whole pool; an event-loop server can stall everything on one thread. It also has nothing to do with which driver you chose — a shipping driver that cannot reach its endpoint is the common case, but a json-file driver on a filesystem that has gone read-only or is out of space blocks in exactly the same way. ## What non-blocking mode does `mode=non-blocking` puts an intermediate in-memory ring buffer between the daemon's read side and the driver. The message is appended to the buffer and the read side continues immediately; a separate consumer drains the buffer into the driver at whatever pace the driver manages. The application's write never waits on the driver. When the driver falls behind far enough that the buffer is full, the buffer **drops log messages**. That is the entire trade: you have chosen to lose logs rather than to lose the workload. It is usually the right choice for a request-serving service and usually the wrong one for an audit trail. `max-buffer-size` sets the buffer's capacity and defaults to `1m`. It is only meaningful with `mode=non-blocking`; in blocking mode nothing is buffered on that path. ``` docker run -d --log-driver=fluentd \ --log-opt mode=non-blocking --log-opt max-buffer-size=4m \ tileserv:2.4 ``` ## Sizing the buffer The buffer must cover *log volume × the stall you expect to ride out*. Take the service's log rate in bytes per second and multiply by the longest driver stall you are willing to survive without loss. A service emitting roughly 180 KB/s of access logs needs about 8.5 MB to cover a 47-second stall; the default `1m` covers under six seconds of it. The reflex to "just make it large" has a cost too — the buffer is daemon memory, held per container, and a large buffer on many containers is real RSS on the host, spent to make loss less likely during an outage in which the logs are usually least valuable. Also be honest about what the buffer is not. It is memory inside `dockerd`. It does not survive a daemon restart, it is not an ordered delivery guarantee, and it is not a substitute for buffering in the collector tier, which is where durability belongs. ## Choosing a default A reasonable posture for a fleet is: request-serving workloads get `mode=non-blocking` with a buffer sized from their measured log rate, so a logging problem can never become a user-facing one; anything whose output is compliance-relevant stays blocking, and its logging path is made reliable enough that blocking never triggers. Whichever you pick, pick it explicitly — the default is blocking, and a team that has never thought about this has implicitly decided that the log backend is a hard dependency of the application. One more reason to decide it deliberately: the option is per container and is frozen when the container is created, like every other log setting. You cannot switch a wedged fleet to non-blocking during the incident that taught you why you wanted it — the change lands only as containers are recreated. That makes this a default to set once, in the same place the fleet's driver and rotation bounds are set, rather than a knob to reach for under pressure.

  • Does mode=non-blocking only matter with remote log drivers?
    No. Back-pressure comes from the driver being slow to accept, whatever the reason. A json-file or local driver on a filesystem that is full, read-only, or backed by a stalled network mount blocks the container's writes exactly like an unreachable log endpoint. Remote drivers make it likelier and more visible, but the local ones are not immune.
  • How do you size max-buffer-size for a service?
    Measure the log rate in bytes per second, decide the longest driver stall you want to ride out without loss, and multiply. A service emitting about 180 KB/s that should survive a 47-second stall needs roughly 8.5 MB, where the 1 MB default covers under six seconds. Then check the total across containers on the host, because the buffer is daemon memory held per container.
  • What evidence would tell you a container was blocked on its log driver rather than on its own work?
    The symptom is application threads stuck in a write to stdout while the process is otherwise idle — no CPU, no database activity, no progress on in-flight requests — and it starts and clears in lockstep with the log destination's availability. Correlating the stall window with the log backend's own outage, and seeing every container on the host affected regardless of workload, points at the shared logging path rather than at any one application.

Blocking mode is a checkout with one till: if the till jams, the queue backs up out of the door and into the street. Non-blocking mode adds a holding pen with a fixed capacity — customers always get in the door, and when the pen is full the overflow is turned away rather than the queue seizing up.

saying these in an interview costs you the question

  • Assumes logging can never affect application latency
  • Thinks non-blocking mode buffers to disk
  • Believes the buffer resends dropped messages later
  • Says only remote drivers can block a container
  • Sets max-buffer-size without setting non-blocking mode
  • Sizes the buffer huge without counting daemon memory

context