skip to content

Why should a containerized service write its logs to stdout instead of a file inside the container?

level: middleimportance: should knowfreq 62%

answer

  1. Which file descriptors does the engine watch
  2. docker logs never reads the filesystem
  3. Ten replicas, ten partial truths
  4. No cron inside a one-process container
  5. The writable layer dies with the container

basics

~20 s

The engine captures the main process's stdout and stderr and hands that stream to the platform's collector; that is also what docker logs shows. A file written inside the container is invisible to that path, per-replica, and discarded when the container is replaced.

solid answer

~50 s

A container's log stream is defined as the standard output and standard error of its main process; the engine captures those two file descriptors and hands them off, which is what `docker logs` reads and what a platform's log collector consumes. Writing to `/var/log/thumbnailer.log` instead puts the data in the container's writable layer, where nothing is watching it. That file is per-container, so ten replicas mean ten partial views of the truth; it is unrotated, because a container runs one process and there is no cron or log rotation daemon in there; and it dies with the container, so the logs from the crash you are investigating disappear with the instance that crashed. Writing to stdout makes the application's job "emit a line" and leaves collection, shipping and retention to the platform, which is the only component that can see all replicas.

code

dockerfile · 5 lines
dockerfile
FROM eclipse-temurin:21-jre
WORKDIR /app
COPY build/libs/thumbnailer.jar /app/thumbnailer.jar
RUN mkdir -p /var/log && ln -sf /dev/stdout /var/log/thumbnailer.log
ENTRYPOINT ["java", "-jar", "/app/thumbnailer.jar"]

go deeper

for a junior

Remember the rule and the reason: the engine captures the main process's stdout and stderr, so anything written to a file inside the container is invisible to docker logs and to the platform's log view.

for a middle

Explain the mechanics -- which descriptors are captured, why a forked helper's output is not, and why a file in the writable layer is per-replica, unrotated and discarded when the container is removed.

for a senior

Demonstrate the failure modes you have actually hit: block buffering hiding the last seconds before a crash, multi-line stack traces shredded into separate records, and a runaway log file consuming the container's ephemeral storage and getting it killed for an unrelated-looking reason.

for a principal

Own where the boundary sits. The image emits one structured record per line and nothing more; collection, shipping and retention belong to the platform. Decide the log schema and the cost of cardinality once, for every service, rather than per team.

### The container log stream is a definition, not a convention When the engine starts a container it does not attach a debugger or tail a directory. It creates pipes for the main process's file descriptors 1 and 2 -- stdout and stderr -- and reads them. Everything downstream is built on that: `docker logs` replays what was captured, and a managed platform's log view is fed from the same stream. So "log to stdout" is not a style preference borrowed from twelve-factor dogma; it is the only channel the surrounding machinery is defined to read. Two consequences follow immediately. First, only the **main** process's descriptors are captured. A background process your entrypoint script forks off, writing to its own file, is outside the stream unless its output is redirected into the parent's. Second, the stream is untyped bytes plus a timestamp -- the engine does not know what a log record is, which is why formatting is entirely the application's problem. ### What actually goes wrong with a file Consider a Spring Boot photo-thumbnail service whose Logback configuration writes a rolling file appender to `/var/log/thumbnailer.log`. Four separate problems appear the moment it runs on a managed platform. **It is invisible.** `docker logs` shows nothing and the platform's log pane is empty, because neither reads the container's filesystem. The first symptom is usually an engineer insisting the app "isn't logging", when in fact it is logging enthusiastically into a place nobody watches. **It is per-replica.** The file lives in that container's writable layer, which is private to that container. Six replicas of the thumbnail service produce six files, each holding a random sixth of the traffic. Answering "what happened to this one upload" now means finding the right replica first, and the replica may already be gone. **It is unrotated.** Inside a container there is normally no cron and no rotation daemon -- the image ships one process, and that is the point. A rolling appender configured inside the app can cap its own file, but anything else that writes (a startup script, a library dumping a heap report) grows without bound and eventually consumes the container's ephemeral storage allowance, at which point the platform kills the container for a disk reason that looks nothing like a logging problem. **It dies with the container.** The writable layer is discarded when the container is removed, and a managed platform removes containers constantly: every deploy, every scale-in, every restart after a crash. The logs you most want are the ones written by the instance that just went away. ### Doing it properly Configure the application's logging framework with a console appender and delete the file appender. For a JVM service that is a one-line change; Spring Boot's default configuration already logs to the console, so the usual fix is removing a `logging.file.name` setting someone added to make local development tidier. For a third-party component you cannot reconfigure, the classic Docker trick is to symlink its log path to the container's stdout at build time -- the official Nginx image does exactly this for its access and error logs. The symlink resolves at runtime through `/dev/stdout` to `/proc/self/fd/1`, so writes land in the captured stream. Its one caveat is that a process which truncates or reopens the file rather than appending to it can break the trick. ### Two details that separate a middle answer from a senior one **Buffering.** When stdout is a pipe rather than a terminal, some runtimes switch from line buffering to block buffering, so output appears in kilobyte-sized bursts or, worse, is lost entirely when the process is killed. Python needs `PYTHONUNBUFFERED=1` or `python -u` for exactly this reason. If your logs arrive minutes late or stop just before a crash, suspect buffering before you suspect the collector. **Multi-line records.** The stream is line-oriented, so a Java stack trace arrives as forty independent lines and a downstream collector will happily index it as forty unrelated records. Structured logging -- one JSON object per line, with the exception rendered into a field -- solves this at the source and is far more reliable than asking every collector to reassemble multi-line events with a regular expression. ### Where the boundary sits The image's contract stops at "emit one well-formed record per line on stdout, with stderr reserved for genuine errors". What happens to the stream afterwards -- which collector consumes it, how much is buffered on the host, how long it is retained -- is the platform's or the engine's configuration, not the image's. That separation is the point: the same image writes the same lines in local development, in a test run and in production, and only the destination changes.

  • Your entrypoint script starts a helper process in the background. Why does its output never appear?
    Only the main process's stdout and stderr are captured. A forked helper writing to its own file or to a descriptor the parent does not own is outside the stream. Redirect it into the parent's stdout, or better, run one process per container so the question does not arise.
  • Logs from a Python-based container arrive in bursts and stop just before a crash. What is the likely cause?
    Block buffering. When stdout is a pipe rather than a terminal the runtime buffers in kilobyte chunks, so records sit unflushed and are lost when the process dies. Setting `PYTHONUNBUFFERED=1` or running `python -u` restores line-by-line flushing. Suspect this before blaming the collector.
  • Should anything go to stderr rather than stdout?
    Keep stderr for genuine error output and diagnostics from the runtime itself, and put application records on stdout. The engine captures both into one stream, so the separation is mostly for humans and for tooling that treats a non-empty stderr as a signal -- but interleaving means ordering between the two is not guaranteed.

saying these in an interview costs you the question

  • Says docker logs reads a file inside the container
  • Adds logrotate and cron to the image
  • Keeps a rolling file appender in production images
  • Assumes logs survive the container being replaced
  • Blames the collector when buffering is the cause
  • Logs multi-line stack traces and expects clean records

context