skip to content

How do you use os.dup and os.dup2 to capture writes to descriptor 1?

level: seniorimportance: should knowfreq 26%

answer

  1. Move the number, not the object
  2. Save a copy before you overwrite
  3. Restore belongs in a finally
  4. A seekable file, never an undrained pipe
  5. Rewind before reading the capture

basics

~20 s

Flush the Python stream, duplicate descriptor 1 with os.dup so you can put it back, os.dup2 your own open file over descriptor 1, run the code, flush again, then os.dup2 the saved duplicate back in a finally block and close it.

solid answer

~40 s

The sequence is save, swap, run, restore. `os.dup(1)` returns a fresh descriptor pointing at the same open file as descriptor 1 — that is your undo handle. `os.dup2(sink.fileno(), 1)` closes whatever descriptor 1 was and makes the number refer to your sink instead, so every writer in the process, compiled code included, now lands there. You flush `sys.stdout` before the swap so nothing already buffered is misfiled, and again before restoring. The restore — `os.dup2(saved, 1)` followed by `os.close(saved)` — belongs in a `finally`, because an exception that skips it leaves the whole process writing into your capture file. `tempfile.TemporaryFile` is the usual sink: seekable, so you can rewind and read it after restoring, unbounded, and self-deleting.

code

python · 16 lines
python
import os, sys, tempfile

sys.stdout.flush()
saved = os.dup(1)
with tempfile.TemporaryFile(mode="w+b") as sink:
    os.dup2(sink.fileno(), 1)
    try:
        print("python-level write")
        os.write(1, b"descriptor-level write\n")
        sys.stdout.flush()
    finally:
        os.dup2(saved, 1)
        os.close(saved)
    sink.seek(0)
    captured = sink.read()
print("captured:", captured)

go deeper

for a junior

Know that descriptor 1 is a number the operating system routes writes through, and that Python can point it somewhere else. You are not expected to write the save-and-restore sequence from memory yet.

for a middle

Be able to write the sequence: flush, os.dup(1) to save, os.dup2 the sink over 1, run, flush, then os.dup2 the saved copy back and close it. Explain why the restore sits in a finally.

for a senior

Demonstrate the operational judgement: choose a seekable temporary file over a pipe and say why, scope the window as narrowly as possible, and describe what a leaked swap looks like in production — output that silently stops with no error raised anywhere.

for a principal

Weigh the technique against alternatives. Descriptor redirection is process-global and hostile to concurrency; argue when to accept that, when to push the noisy component into a separate process instead, and what guardrails keep an ad-hoc capture helper out of a request path.

## Why descriptor level at all Rebinding `sys.stdout` catches Python writers only. When the output you need belongs to a compiled extension, the only interception point is the descriptor number itself, because that is the one thing every writer in the process agrees on. Descriptor-level redirection is the heavier, more global tool, and knowing the exact ritual — and its failure modes — is the interview. ## The five steps **1. Flush first.** `sys.stdout` holds a buffer of its own. Bytes sitting in it were produced before the swap and belong in the old destination; if you swap first they are flushed into your capture file instead. One `sys.stdout.flush()` before the swap removes the ambiguity. **2. Save the descriptor.** `os.dup(1)` allocates the lowest free descriptor number and points it at the same open file description as descriptor 1. It is a second name for the same destination, and it is the only way back — there is no "what was descriptor 1 before" query. The duplicate is non-inheritable, so a child process spawned during the capture will not carry it. **3. Swap.** `os.dup2(sink.fileno(), 1)` atomically closes descriptor 1 if it was open and makes the number 1 refer to the sink's open file. From this instant, `print()`, `os.write(1, ...)` and a compiled writer all land in the sink. Note that `os.dup2` makes the target inheritable by default, so a child process started now inherits the capture too — pass `inheritable=False` if that is not what you want. **4. Run and flush.** Execute the code under capture, then flush `sys.stdout` again so its buffer is emptied into the sink while the sink is still installed. **5. Restore, in a `finally`.** `os.dup2(saved, 1)` puts the original destination back on the number, and `os.close(saved)` releases the duplicate. If an exception escapes the block without running this, descriptor 1 stays pointed at a temporary file for the rest of the process's life — which, on a service, means output vanishing with no error anywhere. Wrap the body in `try`/`finally`, or better, put the whole ritual in a context manager class so the restore is structural. ```python import os, sys, tempfile sys.stdout.flush() saved = os.dup(1) with tempfile.TemporaryFile(mode="w+b") as sink: os.dup2(sink.fileno(), 1) try: print("python-level write") os.write(1, b"descriptor-level write\n") sys.stdout.flush() finally: os.dup2(saved, 1) os.close(saved) sink.seek(0) captured = sink.read() ``` ## Why a temporary file and not a pipe A pipe looks natural — it is what a shell would use — and it is the classic way to hang the process. A pipe has a fixed kernel buffer; once the writer has filled it and nobody is draining the read end, the next write blocks forever. During an in-process capture the only thread that could drain it is the one currently blocked writing to it. `tempfile.TemporaryFile` has no such limit, is seekable so you can `seek(0)` and read the whole capture after restoring, and unlinks itself when closed. Reach for a pipe only when a separate thread owns the read end. ## Consequences worth stating out loud This is **process-global**. There is no per-thread descriptor table: while the swap is installed, every thread, every library and every child process started in that window writes into your sink. In a concurrent service that is a real hazard — unrelated log lines get swallowed into a capture that a different request owns. Confine the window to the narrowest possible call, or serialise captures behind a lock. It also does **not** change `sys.stdout`. The same text wrapper object is still there, still wrapping descriptor 1; only the destination behind the number moved. That is why `sys.__stdout__` is no escape hatch either — it wraps the same number. ## The bugs that show up in review The recurring ones are: overwriting descriptor 1 without saving a duplicate first, so there is no way back; restoring outside a `finally`; calling `os.close(1)` instead of duplicating over it, which leaves the number free for the next `open()` to claim, so an unrelated file silently becomes standard output; forgetting `seek(0)` and reading nothing; and storing the saved descriptor somewhere shared — a module-level list, or a mutable default argument evaluated once at definition time — so that a second, nested or concurrent capture restores the wrong descriptor. Keep the saved value on the stack or on the context-manager instance, one per capture. A concrete shape of that last failure: a fraud-scoring service wraps a compiled scorer's noisy startup diagnostics in a capture helper. The helper stashes the saved descriptor in a mutable default list, so the first call during the roughly 45-second cold start works, and every capture afterwards restores the descriptor saved by the first one. Output stops appearing hours later, and nothing in the traceback points at the helper.

  • Why capture into tempfile.TemporaryFile rather than one end of os.pipe?
    A pipe has a fixed kernel buffer. If the code under capture writes more than fits and no other thread is draining the read end, the write blocks and the capture deadlocks. A temporary file absorbs any volume, is seekable so you can rewind and read it after restoring the descriptor, and removes itself on close. Use a pipe only when a separate reader thread owns the other end.
  • A helper stores the saved descriptor in a mutable default argument. What breaks?
    The default list is created once when the function is defined and shared by every call, so the second capture appends its own saved descriptor next to the first and the restore path uses the wrong one. Descriptor 1 then points at a previous capture file instead of the real destination, and the damage only appears on the second invocation. Keep the saved descriptor in a local or on a context-manager instance.
  • Does swapping descriptor 1 change what sys.stdout is?
    No. sys.stdout is still the same text wrapper around descriptor 1; only the destination behind that number moved, which is precisely why the trick catches compiled writers as well. It also means sys.__stdout__ offers no way back to the terminal during the swap — the only route back is the duplicate you saved.
  • What is the blast radius of this technique in a multithreaded service?
    Total, for the duration. Descriptors are per-process, not per-thread, so every thread, every library and every child process started inside the window writes into your sink. Keep the window as narrow as the call you are capturing, serialise captures behind a lock, and prefer object-level redirection whenever the writers are pure Python.

saying these in an interview costs you the question

  • Overwrites descriptor 1 without saving a duplicate
  • Restores outside a finally, so an exception leaves it swapped
  • Calls os.close(1) instead of duplicating over it
  • Captures into a pipe that nobody is draining
  • Assumes sys.stdout becomes the temporary file object
  • Forgets to seek back to 0 before reading the capture

context