In Ruby, what does IO.pipe return, and why can a read on its reader block forever after the writer has sent everything?
answer
- two connected IO objects
- end of file needs every writer closed
- forked copies of the write end
- block form closes both ends
- a full pipe blocks the writer
basics
~20 sIO.pipe returns [reader, writer], two IO objects joined by an OS pipe. The reader sees end of file only when every copy of the write end is closed, so reader.read hangs while any writer, including one inherited by fork, stays open.
solid answer
~40 s`IO.pipe` creates an operating-system pipe and returns `[read_io, write_io]`; with a block it yields both and closes them when the block ends. Unlike a `StringIO`, these are real `IO` objects with file descriptors, so a forked child can inherit them. `read_io.read` returns only at **end of file**, and a pipe reports end of file only when **every** open copy of the write end is closed. After `fork` both processes hold the write end, so the parent must close its own copy, and the child its read end, or the parent's `read` waits forever. The pipe's kernel buffer is also limited, so a single thread that writes a lot before reading blocks on its own `write`. Ruby puts the write end in sync mode, so writes are not held in a Ruby buffer.
code
ruby · 11 linesreader, writer = IO.pipe
producer = Thread.new do
3.times { |i| writer.puts "line #{i}" }
ensure
writer.close # without this, reader.read never returns
end
puts reader.read # "line 0\nline 1\nline 2\n"
producer.join
reader.closego deeper
Recall that IO.pipe returns a reader and a writer, and that the writer must be closed for the reader's read to finish.
Explain that end of file needs every write descriptor closed, and that the block form closes both ends for you.
Diagnose hangs after fork or from a full pipe, close unused ends in each process, and choose pipe or StringIO for the job.
Prefer higher-level process and queue APIs over raw pipes in application code, keeping raw descriptors to small audited utilities.
## What IO.pipe gives you `IO.pipe` asks the operating system for a **pipe**: a one-way channel with a read end and a write end, backed by a small kernel buffer. Ruby wraps the two descriptors in `IO` objects and returns them as a pair: ```ruby reader, writer = IO.pipe writer.puts "hello" writer.close reader.read # => "hello\n" ``` Facts from the method's documentation and source: - The return value is `[read_io, write_io]`, in that order. - With a block, `IO.pipe { |r, w| ... }` yields both ends, **closes both** when the block exits, and returns the block's value. - Optional encoding arguments set the read end's external and internal encodings, and `binmode: true` opens both ends in binary mode. - The **write end is put in sync mode**, so `writer.puts` goes straight to the kernel rather than into a Ruby-level buffer. - Both ends are real descriptors, so they work with `IO.select`, can be inherited by a forked child, and can be handed to APIs that need a descriptor, all things a `StringIO` cannot do. ## End of file needs every writer closed `reader.read` with no length reads **until end of file**. For a pipe, the kernel reports end of file only when **no process has the write end open**. Data being "all sent" is irrelevant: as long as one descriptor for the write end exists anywhere, the reader assumes more may come and waits. The classic place this bites is `fork`. After a fork the child has copies of every descriptor, so both processes hold both ends. The method's own documentation spells out the discipline: 1. The **parent** closes its write end (`wr.close`) before reading, because it will never write. 2. The **child** closes its read end, writes, then closes its write end. 3. The parent's `rd.read` returns once the child's write end is closed, then the parent reaps the child with `Process.wait`. Skip step 1 and the parent holds a write end itself, so its `read` never sees end of file, even after the child has exited. The same trap appears with threads: if the thread that owns the writer forgets `close`, a reader in another thread blocks forever. ## A full pipe blocks the writer The kernel buffer behind a pipe has a fixed capacity. When it is full, `write` **blocks** until a reader drains some data. That produces a second kind of hang: - a single thread that writes a large payload into the pipe and only then starts reading waits on its own write forever; - the fix is to read and write concurrently, in separate threads or processes, or to use `IO.select` with the non-blocking read and write methods. Once every read end is closed, a write to the pipe fails with `Errno::EPIPE` instead of blocking. ## Reading without waiting for end of file `read` is the only method that insists on end of file. The others return as soon as enough data has arrived: - `reader.gets` returns one line as soon as a newline is in the pipe, and `nil` at end of file, so a `while (line = reader.gets)` loop processes output as it is produced. - `reader.readpartial(4096)` returns whatever bytes are available, blocking only when none are, and raises `EOFError` at the end. - `reader.read_nonblock(4096)` never blocks: with nothing available it raises an exception that includes `IO::WaitReadable`, or returns `:wait_readable` when called with `exception: false`, which pairs with `IO.select`. Line-by-line reading does not remove the end-of-file rule; it only means you see data sooner. The loop still ends only when every write end is closed. ## Pipe versus StringIO | Need | `IO.pipe` | `StringIO` | |---|---|---| | In-memory buffer inside one thread | awkward: capacity limits, needs a second thread | natural | | Hand data to a forked child | yes, the child inherits the descriptors | no descriptor to inherit | | Wait for data with `IO.select` | yes | no | | Test code that only calls `puts` or `gets` | works, but heavier | simplest choice | Reach for `IO.pipe` when something needs a **real descriptor** or a **producer and consumer running concurrently**. For capturing text inside one process, `StringIO` is simpler and has no capacity or end-of-file rules to get wrong. ## Checklist - Close the end you do not use in each process or thread, immediately. - Close the write end when you finish writing, so the reader sees end of file. - Never fill a pipe from the same thread that is supposed to drain it. - Prefer the block form when both ends live in one scope, so they are closed even when an exception escapes.
- Why does the parent's read hang even after the child process has exited?The parent still holds its own copy of the write end, inherited from before the fork. The kernel reports end of file only when every write descriptor is closed, and the parent's copy is still open, so `read` keeps waiting. Closing `wr` in the parent before reading fixes it.
- What happens when a thread writes more than the pipe can hold before anyone reads?The write blocks once the kernel's pipe buffer is full and stays blocked until a reader drains data. If the same thread was going to read afterwards, it deadlocks on its own write. Reading and writing must run concurrently.
saying these in an interview costs you the question
- IO.pipe returns the writer first and the reader second
- reader.read returns as soon as the writer has sent all its data
- A pipe can buffer any amount of data, so writes never block
- Closing the write end in the parent after fork is optional tidiness
- StringIO can replace IO.pipe for passing data to a forked child