skip to content

Why does output written before os.fork() get printed twice?

level: juniorimportance: should knowfreq 40%

answer

  1. The copy includes more than your code
  2. Nothing was printed twice; something pending was
  3. Userspace buffering, not the terminal
  4. Empty it before you duplicate it
  5. sys.stdout.flush() before os.fork()

basics

~10 s

os.fork() copies the whole process, including text still sitting in stdout's userspace buffer. Both copies flush that same pending text later, so it prints twice. Call sys.stdout.flush() immediately before forking.

solid answer

~40 s

`sys.stdout` is a buffered writer: when it is not attached to a terminal it collects up to a few kilobytes in a userspace buffer before writing anything to the file descriptor. `os.fork()` duplicates the address space, so the child gets its own copy of that buffer with the same unwritten bytes in it. Each process then flushes its copy when it exits, and the same text reaches the file description twice. The fix is to empty the buffer before you fork -- `sys.stdout.flush()` (and `sys.stderr.flush()`, plus any other buffered writer you have open) as the last statement before `os.fork()`. Exiting the child with `os._exit()` also avoids it, because `os._exit()` discards the buffer instead of flushing it, but then the child must flush its own output explicitly.

code

python · 8 lines
python
import os, sys

sys.stdout.write("balances: ")   # no newline: still sitting in the buffer
if os.fork() == 0:
    sys.exit(0)                  # the child flushes the inherited copy too
os.wait()
sys.stdout.write("done\n")
# prints: balances: balances: done

go deeper

for a junior

Recall that os.fork() copies the running process, including whatever text is still waiting in stdout's buffer, and that calling sys.stdout.flush() right before the fork is the fix.

for a middle

Explain the buffering mode that makes it happen: a terminal is line-buffered so the buffer is usually empty, a pipe or file is block-buffered so pending bytes survive the copy and each process flushes them.

for a senior

Show how you would catch this in a service whose logs go through a pipe: flush before every fork, hard-exit children with os._exit(), and treat every buffered writer open at fork time as a suspect, not just stdout.

for a principal

Own the rule for the codebase -- no process forks with dirty buffers. Decide whether fork sites go through one helper that flushes and hard-exits, or whether the job should use a higher-level process API and not call os.fork() at all.

### What `os.fork()` actually copies `os.fork()` creates a child process whose memory is a copy of the parent's. That copy is not limited to your data structures: it includes every object CPython's runtime holds, and among those are the `io` objects behind `sys.stdout` and `sys.stderr`. A text stream like `sys.stdout` is a `TextIOWrapper` wrapped around a buffered binary writer, and that buffered writer owns a chunk of plain memory holding bytes you have "written" but that have not yet been handed to the operating system. So at the moment of the fork there are two independent copies of that buffer, both holding the same unwritten bytes, and both attached (through the inherited file descriptor) to the same open file description. Whatever each process does with its copy later lands in the same place. ### Why the bytes are still pending Python chooses a buffering mode for the standard streams based on what they are connected to. When `sys.stdout` is a terminal it is line-buffered, so every newline pushes the line out and the buffer is usually empty. When it is redirected to a file, a pipe, a log collector or a supervisor -- which is exactly what happens in production, in CI and under a process manager -- it is block-buffered, and text can sit unwritten until the buffer fills or the interpreter exits. `sys.stderr` has been line-buffered by default since Python 3.9 even when it is not a terminal, which is why the duplication is usually seen on stdout first. This is the reason the bug has the reputation of being intermittent: run the script by hand in a terminal and each `print()` flushes at its newline, so there is nothing pending to duplicate; run it under a pipeline and the same code prints its preamble once per process. ### What each process does with its copy If the child leaves through the ordinary exit path -- falling off the end of the script, raising `SystemExit` via `sys.exit()`, or dying on an uncaught exception -- the interpreter shuts down normally in the child, and part of that shutdown is flushing the standard streams. The child writes the inherited pending bytes. Later the parent exits and flushes its own copy of the same bytes. One `print()` in the parent, two lines in the output. Nothing about this is specific to `print()`. Any buffered writer that is open with a non-empty buffer at fork time behaves the same way: a log file opened for writing, an `io.BufferedWriter` over a socket-backed file object, a CSV report being assembled. If a fork happens while those buffers are dirty, every child that exits cleanly re-emits them. ### The fixes, in order of preference **Flush before you fork.** The direct fix, and the one that keeps the child's own behaviour unchanged: ```python sys.stdout.flush() sys.stderr.flush() pid = os.fork() ``` After the flush the buffers are empty, so copying them copies nothing. **Hard-exit the child.** `os._exit(status)` calls the underlying `_exit` system call immediately: it does not run interpreter shutdown, so it does not flush anything. A child that ends in `os._exit()` therefore cannot duplicate inherited output. The tradeoff is that the child's *own* output is discarded too, so a child that prints anything must call `sys.stdout.flush()` itself before hard-exiting. In practice you want both: flush before the fork, and hard-exit the child. **Turn buffering off wholesale.** Running the interpreter with `-u`, or setting the `PYTHONUNBUFFERED` environment variable, makes the standard streams unbuffered (stdout becomes write-through), so there is never anything pending. `sys.stdout.reconfigure(line_buffering=True)` is the narrower in-process version. These are blunt instruments -- they cost a syscall per write -- but for a job whose output is a modest log stream they are often the right blunt instrument anyway, and they make the ordering of parent and child lines predictable as a bonus. ### The lookalike bug One symptom, two causes. Duplicated output can also mean the child never left the fork branch at all and is re-running the parent's code as a second copy of the program -- typically because `sys.exit()` in the child was swallowed by an enclosing `except` clause. Check which one you have: a buffer-duplication bug repeats only text written *before* the fork, and repeats it verbatim; a runaway child repeats work done *after* the fork, and usually shows a second process in the process table doing real work. ### Version note None of this changed between Python 3.10 and 3.14. `os.fork()` remains a thin wrapper over the system call with the same buffer-copying consequences it has always had.

  • Why does the duplication disappear when you run the same script interactively in a terminal?
    Buffering mode depends on what the stream is connected to. On a terminal `sys.stdout` is line-buffered, so each newline empties the buffer and there is nothing pending at fork time. Redirected to a file or a pipe it is block-buffered and can hold several kilobytes unwritten, which is exactly the state that gets copied. The bug is therefore invisible by hand and reproducible under a pipeline or a process manager.
  • Does flushing stdout cover every stream that can duplicate output after a fork?
    No. Any buffered writer that is open with a dirty buffer at fork time duplicates: an application log file, a report file being assembled, an `io.BufferedWriter` over another descriptor. `sys.stderr` is line-buffered by default so it rarely bites, but the rule is about all of them -- flush everything you hold, or make the child hard-exit with `os._exit()` so it never flushes anything it inherited.

A pending buffer is a letter still sitting in your outbox. os.fork() photocopies the entire desk, outbox included, so two people each post the same letter.

saying these in an interview costs you the question

  • Claims os.fork() re-executes the parent's earlier print statements
  • Thinks the terminal, not Python, duplicates the text
  • Says print() is not atomic so the write got retried
  • Believes os.fork() gives the child fresh empty stdio buffers
  • Suggests os.sync(), which flushes kernel buffers, not Python's
  • Assumes it cannot happen because it never reproduces in a terminal

context