skip to content

Why can contextlib.redirect_stdout fail to capture the output you expected?

level: seniorimportance: should knowfreq 28%

answer

  1. It swaps a name, nothing more
  2. Anything that saved the old stream misses it
  3. One process, not one thread
  4. The OS never sees sys.stdout
  5. Children and C code write to descriptor 1

basics

~20 s

It only rebinds the name sys.stdout, process-wide, for the block's duration. Writes through a saved stream reference, straight to descriptor 1 from C code, or from a child process escape it; other threads' prints get captured.

solid answer

~40 s

`contextlib.redirect_stdout(target)` is a small manager: `__enter__` saves the current `sys.stdout`, assigns the target in its place and returns the target; `__exit__` restores the saved value. Nothing is teed or duplicated. Three consequences bite in production. It is a **name rebinding**, so anything holding the stream object it was given earlier — a `logging.StreamHandler` configured at startup, a cached `write` method — keeps writing to the real console. It is **process-global**, not thread-local or task-local, so concurrent threads print into your buffer and their output vanishes from the console. And it operates at the **Python level**, so an extension module writing straight to file descriptor 1, or a child process spawned with `subprocess.run`, is entirely unaffected. Capturing those needs `os.dup2` at the descriptor level. `redirect_stderr` is the same class pointed at `sys.stderr`.

code

python · 9 lines
python
import contextlib
import io

buf = io.StringIO()
with contextlib.redirect_stdout(buf):
    print("captured")
    print("also captured")

print("buffer holds:", repr(buf.getvalue()))

go deeper

for a junior

Know the happy path: it points sys.stdout at your object for the duration of the block, so print goes into a buffer you can read afterwards with getvalue().

for a middle

Explain the mechanism — save, rebind, restore — and derive from it that anything caching the old stream, or writing below the Python level, is unaffected.

for a senior

Be ready to diagnose real symptoms: output missing from the console because another thread's print was captured, an empty buffer because the work happened in a child process, and know the descriptor-level fix with os.dup2.

for a principal

Own the design position: global stream swapping is a test-harness tool, not a library technique. Push for injected streams or logging in code you control, and confine redirection to boundaries you can prove are single-threaded.

### The mechanism, in full `contextlib.redirect_stdout(new_target)` does almost nothing, and knowing exactly how little is the answer to the question. On entry it reads the current value of `sys.stdout`, pushes it onto an internal stack of saved values, assigns `new_target` to `sys.stdout`, and returns `new_target` so it can be bound with `as`. On exit it pops the saved value back into `sys.stdout`. `redirect_stderr` is the identical class aimed at `sys.stderr`. That is a **name rebinding on the `sys` module**, and every limitation follows from it. ### Limitation 1: only lookups through `sys.stdout` are affected `print()` looks up `sys.stdout` at call time, which is why the common case works. Anything that resolved the stream *earlier* did not resolve a name, it captured an object, and it keeps writing to that object: ```python handler = logging.StreamHandler() # grabs sys.stderr now ``` The classic report is "my logger still prints to the console inside the redirect". It does, and correctly so: the handler holds the original stream. Only handlers created inside the block, or explicitly pointed at the new target, follow the redirect. The same applies to any code that did `out = sys.stdout` at import time, and to `print(..., file=saved)`. ### Limitation 2: it is process-global `sys.stdout` is one module attribute for the whole process. It is not thread-local, not task-local, and not interpreter-scoped in any way that helps you. So while one thread is inside a `redirect_stdout` block, **every** thread's `print` lands in that thread's buffer, and the output that was supposed to reach the console disappears into it. Worse, when the block exits it restores whatever it saved — if two threads redirect concurrently, the restores can interleave and leave `sys.stdout` pointing at a buffer that is no longer in use. The rule is simple: use it in single-threaded code, or in a process you fully control for the duration. The same reasoning applies under `asyncio`: an `await` inside the block yields to other tasks that will print into your buffer. Nesting in a *single* thread, by contrast, is fine — the saved values form a stack, so nested and repeated use of the same instance restore correctly. ### Limitation 3: it is a Python-level construct The operating system knows nothing about `sys.stdout`. Output written directly to file descriptor 1 bypasses it entirely, and there are two big sources of that: - **Extension modules.** A compiled extension that writes with C-level `printf` or `fwrite` goes to the descriptor, not through the Python object. - **Child processes.** `subprocess.run([...])` hands the child the parent's descriptors. The child prints to the terminal while your buffer stays empty. Capturing a child's output is done with the `subprocess` call's own capture arguments, not with redirection of the parent's `sys.stdout`. To catch everything the process emits, you must redirect at the descriptor level instead: save a duplicate of descriptor 1 with `os.dup`, point descriptor 1 at your file with `os.dup2`, run the code, then `os.dup2` the saved copy back. That is heavier, affects children and extensions too, and is genuinely global — which is why it belongs in a test harness, not in library code. ### What it is good for Given all that, it remains the right tool in a narrow band: capturing what a function prints, in single-threaded code, when you cannot change that function. Tests that assert on printed output, CLI code that needs to tee a help dump into a string, notebook-style capture of a chatty helper. `io.StringIO` is the usual target; a real file works equally well. The design lesson underneath is worth stating in an interview: needing `redirect_stdout` is usually a signal that a function should have taken a stream parameter or used the logging machinery in the first place. `print(msg, file=out)` with `out` defaulting to `sys.stdout` is testable without touching global state at all. Redirection is what you reach for when the code is not yours to change. ### Sharp edges to mention - The target must be a file-like object with the write methods the code will actually call; passing something that only implements `write` breaks a caller that also flushes. - The buffer is only readable after you call `getvalue()` on it — inside the block, printing your own progress goes into the buffer too, which is a surprisingly common self-inflicted confusion. - On exit it restores the *saved* value, not `sys.__stdout__`, so it composes correctly with an outer redirection but will faithfully restore an already-broken stream if one was in place. - Free-threaded builds change nothing here: `sys.stdout` is still one name shared by all threads, and the race is if anything easier to hit. ### What to say in an interview "It swaps the `sys.stdout` name for the duration of the block and restores it after — that is all. So it misses anything that cached the old stream, anything writing to descriptor 1 from C or a child process, and it is global rather than per-thread, so concurrent prints land in my buffer. I use it single-threaded for capture, and reach for descriptor-level redirection when I need the real thing."

  • How would you capture output that a compiled extension writes straight to file descriptor 1?
    Redirect at the OS level rather than the Python level: save a duplicate of descriptor 1 with `os.dup`, point descriptor 1 at a temporary file with `os.dup2`, run the code, then `os.dup2` the saved descriptor back and read the file. That catches everything the process writes, extensions and child processes included, at the price of being even more global than the Python-level swap.
  • Why does an already-configured logger still print to the console inside a redirect_stdout block?
    A stream handler stores the stream object it was given when logging was configured — usually the original `sys.stderr` or `sys.stdout` — so rebinding the name afterwards changes nothing for it. Only handlers created inside the block, or ones explicitly reassigned to the new target, follow the redirection.
  • Is redirect_stdout safe to nest, and is it safe with threads?
    Nesting in one thread is safe: it keeps a stack of saved values, so inner blocks restore to the outer target correctly and one instance can be reused. Threads are the opposite — the swap is a process-wide name rebinding, so any thread printing during the block writes into your target, and concurrent restores can leave sys.stdout pointing at a dead buffer.
  • What would you change in the code under test so redirection is unnecessary?
    Give the function a stream parameter defaulting to None and resolved to `sys.stdout` at call time, or route the messages through logging. Then a test passes its own buffer and asserts on it with no global state touched. Redirection is the fallback for code you cannot change.

It is redirecting one person's mail at the front desk, not at the post office: letters addressed through the desk get moved, but anyone who already wrote down the old address, and every courier bypassing the desk, still delivers to the original.

saying these in an interview costs you the question

  • Thinks it captures a subprocess's output
  • Assumes the redirection is thread-local or task-local
  • Believes it tees output to both destinations
  • Expects it to catch C-level writes to descriptor 1
  • Thinks it redirects stderr as well as stdout
  • Forgets handlers configured earlier keep the old stream

context