Compare the reactor pattern — the operating system tells you a handle is ready and you then perform the I/O yourself — with the proactor pattern, where you submit the I/O and are told when it completed. What changes for buffer ownership, error handling, and portability?
answer
- readiness = "you may act"; completion = "it is done"
- proactor pins a buffer per in-flight op
- errors on your call vs on the completion record
- files: reactor can't, proactor can
- proactor APIs often emulated over readiness
basics
~20 sReactor: demultiplex readiness, then your loop does the transfer with your buffer, allocated on demand; errors surface on your call. Proactor: submit operation plus buffer, kernel transfers, completion event carries bytes and error. Proactor pins buffers per in-flight operation and batches better; reactor is more portable.
solid answer
~60 sBoth are demultiplexer patterns: one loop waits on many operations and dispatches handlers. They differ in *what the event means*. **Reactor** — the event says "this handle would not block now". The handler performs the read or write, so the copy runs on the loop thread. Buffers can be allocated only when data is known to be waiting, so buffer memory tracks active connections. Errors arrive from the call the handler makes. It maps onto every readiness interface, so it ports easily. **Proactor** — the event says "the operation you submitted is finished". You supply the buffer at submission and must not touch it until completion; buffer memory therefore tracks *outstanding operations*, including idle sockets with a read posted. The kernel performs the transfer, errors ride on the completion record, and submissions batch well, so system calls per operation can approach zero. Practically: proactor covers regular files properly, reactor does not. Most cross-platform runtimes implement a proactor-shaped API on top of whichever kernel mechanism exists, emulating the missing half.
code
text · 11 linesREACTOR
register(sock, READABLE)
on_ready(sock):
n = read(sock, my_buffer) # copy happens here, my thread
parse(my_buffer, n)
PROACTOR
submit_read(sock, buf_A, op_id=7) # buf_A now kernel-owned
on_completion(ev):
if ev.op_id == 7:
parse(buf_A, ev.bytes) # already copied; buf_A mine againgo deeper
Distinguish the two events: "ready, go ahead" versus "already done". That alone is a good junior answer.
Add who performs the copy and who owns the buffer, and name readiness queues versus completion ports as the typical mechanisms.
Cover buffer-memory scaling, error and cancellation semantics, the regular-file gap, and that most portable libraries emulate proactor over reactor.
Reason about it as a substrate choice driven by workload shape — connection-heavy versus operation-heavy versus file-heavy — including batching economics and the memory ceiling posted reads impose.
## The two patterns, precisely Both patterns solve the same problem — one thread, many concurrent I/O operations, no thread parked per operation — and both are built from the same three parts: a *demultiplexer* that waits for many things at once, a *dispatcher* that maps an event back to the code that cares, and *handlers* holding per-connection state. The difference is the semantics of the event. **Reactor** (readiness-based). Handles are put in non-blocking mode and registered with a readiness demultiplexer. The loop waits; the kernel says which handles are now readable or writable; the dispatcher calls the handler; the handler performs the actual read or write, which is guaranteed not to park because readiness was just reported. The data copy happens inside your loop, on your thread, with your buffer. **Proactor** (completion-based). The application initiates the operation: "read up to N bytes from this handle into this buffer". The call returns immediately. The kernel performs the transfer asynchronously. The demultiplexer waits for *completions*; each completion carries the operation identity, the byte count, and any error. The dispatcher hands that to the handler, which now works with data that is already in memory. ## Buffer ownership This is the sharpest practical difference. - Reactor: you learn there is data, *then* you choose a buffer. You can use one shared scratch buffer per loop thread for reads and copy out only what you keep. Ten thousand idle connections need almost no buffer memory. - Proactor: a buffer must be committed at submission and belongs to the kernel until completion. If you post a read on every idle connection so you learn promptly when data arrives, you have pinned one buffer per connection. Ten thousand idle connections with a 16 KB posted read is 160 MB of untouchable memory. Real proactor systems fight this with small "zero-byte reads" that only signal arrival, with shared buffer pools registered with the kernel, or with buffer-selection features where the kernel picks a buffer from a pool at completion time. The corollary bug class is unique to proactor: freeing, reusing, or reading a buffer that still has an operation in flight, and forgetting that cancellation is *asynchronous* — a cancelled operation may still complete, so the buffer is not yours again until you see its completion event. ## Error handling and cancellation Reactor errors are ordinary: your read or write returns a failure, in the handler, on the stack you already have. Proactor errors are delivered as a field on a completion for an operation you submitted possibly long ago, so handlers must carry enough context to interpret them, and every operation needs a durable identity. Timeouts differ too: in a reactor you simply stop waiting; in a proactor you must request cancellation and then still reap the completion. ## Where the copy runs, and throughput In a reactor every event costs at least one wait plus one read call, and the copy consumes loop-thread cycles. Under very high event rates those system calls dominate. Proactor implementations can amortise both: batch many submissions into one call, reap many completions at once, keep buffers and handles pre-registered with the kernel to skip per-call validation, and let hardware move bytes. That is precisely the design of modern submission/completion ring interfaces. ## Regular files Readiness models were designed for sockets and pipes. A regular file is always reported ready, yet reading it can still park the thread on device I/O. So a pure reactor cannot do asynchronous file I/O at all — real systems shove file work onto a worker thread pool. Completion models handle files natively, which is a major reason they exist and a strong argument for them in storage-heavy services. ## Portability and emulation Readiness demultiplexers exist nearly everywhere; the canonical completion-port model came from one family of systems, with ring-based completion interfaces arriving on others much later. Consequently: - A proactor API can be *emulated* over a reactor: wait for readiness, do the read into the caller's buffer yourself, then synthesise a completion event. Cross-platform networking libraries have done exactly this for two decades, which is why their APIs look completion-shaped everywhere. - The reverse emulation — reactor over a completion port — is awkward, because you must post an operation to learn about readiness, which brings the buffer-pinning problem back. So the pattern you *program against* is often not the mechanism the kernel provides. When someone says "we use a proactor", the honest follow-up is whether that is a native completion interface or an emulation over readiness, because the buffer and system-call economics differ entirely. ## Choosing Prefer a reactor when connections vastly outnumber active transfers, memory is tight, and portability matters. Prefer a proactor when you do heavy file I/O, when per-operation system-call cost is your ceiling, or when the platform's native model is completion-based and emulating readiness would fight the kernel. In both cases the handler discipline is identical and non-negotiable: never block, never run long, inside the loop.
- Why does a proactor design use more memory for 50,000 idle connections than a reactor design?To be notified promptly, a proactor posts a read on each idle connection, and every posted read pins a buffer the kernel owns until completion — so buffer memory scales with outstanding operations rather than with active ones. A reactor allocates a buffer only after readiness is reported, so idle connections cost only their state record. Mitigations exist: zero-byte reads that signal arrival without a buffer, or kernel-side buffer pools where a buffer is chosen at completion time.
- Your networking library exposes a completion-style API on every platform. What should you check before assuming completion-model performance?Check whether the platform has a native completion interface or the library is emulating one over a readiness mechanism. Emulation still performs the copy on the loop thread and still pays a system call per operation, so the batching and file-I/O advantages of a true completion model are absent — and file operations are probably being handed to a hidden worker pool. The API shape tells you nothing about the underlying economics.
A reactor is a receptionist who tells you a visitor has arrived so you go and collect them; a proactor is a courier service you dispatch, which reports back once the delivery is already on your desk.
saying these in an interview costs you the question
- Treating reactor and proactor as synonyms for "async" versus "sync".
- Reusing or freeing a submitted buffer before its completion event arrives.
- Assuming cancellation in a completion model takes effect immediately.
- Claiming a reactor can perform true asynchronous regular-file I/O.
- Assuming a completion-shaped API implies a completion-based kernel mechanism.