skip to content

Why does contextlib.redirect_stdout miss output from a compiled extension?

level: middleimportance: must knowfreq 45%

answer

  1. Two stacked layers share one name
  2. The name versus the number
  3. Compiled code never reads a Python attribute
  4. Descriptor 1 is the real destination
  5. os.dup2 moves the number itself

basics

~20 s

contextlib.redirect_stdout only rebinds the Python object sys.stdout. Compiled code writes straight to file descriptor 1 and never consults sys.stdout, so its output bypasses the capture; redirecting the descriptor itself with os.dup2 is what catches it.

solid answer

~40 s

`contextlib.redirect_stdout(target)` is a small context manager: on entry it saves `sys.stdout` and assigns your object to it, on exit it puts the old one back. Everything reaching output through `print()` or `sys.stdout.write()` therefore lands in your target, because those resolve the attribute at call time. A compiled extension does not resolve anything Python-side — it issues a C-level write on file descriptor 1, which the `sys` module has no say over. The same holds for a child process, which inherits the descriptor, and for any module that captured the old stream object earlier. To intercept those you must change what descriptor 1 points at: save it with `os.dup`, overwrite it with `os.dup2` pointing at a real file such as a `tempfile.TemporaryFile`, and restore it afterwards. Object-level and descriptor-level redirection are different layers.

code

python · 7 lines
python
import contextlib, io, os

buf = io.StringIO()
with contextlib.redirect_stdout(buf):
    print("goes into the buffer")
    os.write(1, b"goes to the terminal\n")
print("captured:", buf.getvalue().strip())

go deeper

for a junior

Recall that print() sends text to whatever object sys.stdout currently names, and that contextlib.redirect_stdout swaps that object for the duration of a with block. Knowing that a file descriptor is a separate, lower-level thing is enough at this level.

for a middle

Explain the mechanism in both directions: redirect_stdout saves and reassigns one module attribute, while a compiled writer issues a syscall on descriptor 1 that never touches Python state. Name os.dup2 as the layer that catches the second kind.

for a senior

Show the diagnosis, not just the fact. Be ready to say how you would confirm a capture is partial, why the missing lines are not a buffering problem, and what a descriptor-level capture costs you in a process where other threads are also writing.

for a principal

Own the policy question: when a service needs to quarantine a noisy dependency's output, argue for fixing it at the source, at the descriptor, or at the process boundary, and weigh the process-global blast radius of descriptor redirection against the partial coverage of the object-level tool.

## Two layers wearing the same name The word "stdout" names two stacked things in a Python process, and this question is entirely about the gap between them. At the bottom is **file descriptor 1**, a small integer the kernel associates with an open file, pipe or terminal. Any code in the process — Python, C, assembly — can issue a write syscall against that number, and the kernel routes the bytes wherever the descriptor currently points. The interpreter has no involvement at all. On top sits **`sys.stdout`**, an ordinary Python object: a text wrapper that encodes `str` to bytes, buffers them, and eventually calls write on descriptor 1. `print()` does not talk to the kernel. It looks up the name `stdout` on the `sys` module at call time and calls that object's `write` method. ## What redirect_stdout actually does `contextlib.redirect_stdout` works purely on the upper layer. Its entire mechanism is: remember the current value of the `sys.stdout` attribute, assign the target you passed in, and on `__exit__` assign the remembered value back. That is it — no descriptors, no syscalls, no operating-system state. `contextlib.redirect_stderr` is the identical mechanism for `sys.stderr`. Because `print()` re-resolves `sys.stdout` on every call, the substitution is invisible to normal Python code, and an `io.StringIO` target collects everything cleanly. That is why the tool feels like it captures "the output" — for pure-Python writers, it does. ## The writers it cannot see Several classes of writer never consult `sys.stdout`, and each of them escapes the capture: 1. **Compiled code.** A C extension that prints diagnostics calls the C runtime's write on descriptor 1 directly. Nothing in that path reads a Python attribute. 2. **A child process.** It inherits descriptor 1 from the parent; it has its own interpreter state, or none at all, and certainly not your `StringIO`. 3. **Early-bound references.** A module that did `from sys import stdout` at import time, or a logging handler configured before the block, holds a direct reference to the *old* object. Rebinding the module attribute does not reach through an existing reference. 4. **Anything writing to the descriptor deliberately**, such as `os.write(1, ...)` — useful as a stand-in for a compiled writer when you want to reproduce the effect. The symptom in practice is a silent partial capture: your buffer holds the Python-side lines, the extension's diagnostics still appear on the terminal or in the process's log file, and nothing raises. A common wrong diagnosis is "the output was buffered and lost"; the bytes were never routed anywhere near your object. ```python import contextlib, io, os buf = io.StringIO() with contextlib.redirect_stdout(buf): print("python-level write") os.write(1, b"descriptor-level write\n") # only the first line is in buf ``` ## The fix is one layer down To catch every writer you must change what descriptor 1 *is*, not what `sys.stdout` *names*. The standard sequence is: flush the Python-side stream so nothing already buffered escapes into the wrong place, duplicate descriptor 1 with `os.dup` so you can put it back, `os.dup2` a real file over descriptor 1, run the code, flush again, then `os.dup2` the saved duplicate back and close it. `tempfile.TemporaryFile` is the usual sink because it is seekable, unbounded, and unlinks itself. After that swap, `print()` and the compiled extension both land in the file — not because `sys.stdout` changed (it did not; it still wraps descriptor 1) but because the number now points somewhere else. ## Choosing between them Object-level redirection is cheap, exception-safe, portable to Windows, and confined to the interpreter, so prefer it when the writers are Python. Descriptor-level redirection is heavier and genuinely process-global — it affects every thread and every library at once, and there is no per-thread version of it — but it is the only thing that sees compiled and inherited writers. One further caveat about the object-level tool: `sys.stdout` is a module attribute shared by the whole process, so `contextlib.redirect_stdout` is *not* thread-safe. A background thread printing during the block writes into your buffer as well. That is a property of the layer, not a bug in the tool, and it is another reason to reach for a purpose-built capture rather than redirection inside a concurrent service. Finally, note that `sys.__stdout__` is untouched by all of this: it still refers to the startup object, which is exactly why it is the conventional way back if a library reassigns `sys.stdout` and forgets to restore it.

  • Is contextlib.redirect_stdout safe to use in a multithreaded process?
    No. It assigns a module attribute that the whole process shares, so any other thread that calls print during the block writes into your target too, and the restore on exit can race with a concurrent redirect. It is fine in single-threaded scripts and tests; in a concurrent service prefer passing an explicit stream parameter to the code you want to capture.
  • Why can pure-Python output still escape a redirect_stdout block?
    Because rebinding an attribute does not reach through references that already exist. A module that did `from sys import stdout` at import time, or a logging handler that was given the stream object when it was configured, holds the original object directly and keeps writing to it. Only code that resolves `sys.stdout` at call time follows the substitution.
  • How would you demonstrate the gap without a compiled extension to hand?
    Call `os.write(1, b"...")` inside the block. That is the same syscall a compiled writer makes, it needs nothing outside the standard library, and it shows the bytes reaching the terminal while `print()` output lands in your buffer — a two-line reproduction of the whole failure.

saying these in an interview costs you the question

  • Claims redirect_stdout changes file descriptor 1
  • Thinks print writes to the descriptor directly
  • Says a StringIO target captures every writer in the process
  • Assumes a child process inherits the sys.stdout object
  • Believes sys.__stdout__ is updated by the redirect
  • Blames buffering for output that was never routed there

context