skip to content

What are the common pitfalls when writing a Selector-based server, and how do you avoid them?

level: seniorimportance: should knowfreq 40%

answer

  1. it.remove() every handled key
  2. OP_WRITE only on partial write; clear when drained
  3. read(): 0=nothing, -1=closed; writes can be partial
  4. Don't block the loop — offload slow work
  5. Buffer cycle: read -> flip -> consume -> compact/clear
  6. epoll spin bug -> rebuild selector (or use Netty)

basics

~20 s

Main pitfalls: forgetting to remove handled keys from selectedKeys() (causes reprocessing), leaving OP_WRITE always on (busy spin), ignoring partial reads/writes, blocking inside the loop, and not handling -1 from read() (peer closed). Fix each by following the correct NIO idioms.

solid answer

~50 s

The classic Selector bugs are: (1) not calling iterator.remove() on each handled key, so the selector keeps redelivering it; (2) keeping OP_WRITE in the interest set permanently — sockets are almost always writable, so select() spins — instead enable OP_WRITE only when a write is partial and clear it when drained; (3) treating read()/write() as all-or-nothing, when reads return partial data (and -1 means the peer closed: you must cancel/close) and writes may not consume the whole buffer; (4) doing blocking or slow work on the event-loop thread, stalling every connection — offload to a worker pool; (5) ByteBuffer misuse — forgetting flip() before writing out what you read, or clear()/compact() before reading more; (6) the JDK epoll selector-spin bug where select() returns 0 in a tight loop, mitigated by detecting the spin and rebuilding the selector (or just using Netty). Also: handle CancelledKeyException and close channels to release fds.

code

java · 12 lines
java
// Correct read handling with partial-read / close awareness
SocketChannel ch = (SocketChannel) key.channel();
ByteBuffer buf = (ByteBuffer) key.attachment();   // per-connection state
int n = ch.read(buf);
if (n == -1) {            // peer closed
    key.cancel();
    ch.close();           // free the file descriptor
} else if (n > 0) {
    buf.flip();           // switch to read mode
    process(buf);         // may consume only part of a message
    buf.compact();        // keep undrained bytes for next read
}

go deeper

for a junior

Recognizes the must-remove-keys rule and that read() can return -1 when the peer closes.

for a middle

Handles partial reads/writes, the ByteBuffer flip/compact cycle, and avoids permanent OP_WRITE.

for a senior

Knows to offload blocking work, manage write interest on demand, guard cancelled keys, close channels to free fds, and the epoll-spin mitigation.

for a principal

Designs the loop's threading/back-pressure model, cross-thread mutation via wakeup(), fd-budget and graceful shutdown, and argues for framework adoption with the trade-offs.

Writing a correct `Selector` loop by hand is famously tricky. Here are the pitfalls interviewers probe, each with the underlying mechanism and the fix. ## 1. Not removing handled keys `selector.select()` *adds* ready keys to `selector.selectedKeys()` but **never removes** them. If you iterate without `it.remove()`, the same key reappears next cycle and you reprocess a possibly-stale event — leading to spinning or acting on a no-longer-ready channel. **Fix:** always `it.remove()` right after taking the key: ```java while (it.hasNext()) { SelectionKey k = it.next(); it.remove(); /* handle */ } ``` ## 2. Permanent OP_WRITE interest (busy spin) A connected socket's send buffer almost always has room, so it is **writable nearly all the time**. If `OP_WRITE` stays in the interest set, `select()` returns immediately every cycle and you burn a CPU core. **Fix (demand-driven writes):** keep interest at `OP_READ`; attempt `write()` directly; only if it's a **partial write** add `OP_WRITE`; once the buffer is fully drained, remove `OP_WRITE` again. ## 3. Ignoring partial reads and writes Non-blocking I/O does **not** guarantee a whole message moves in one call. - `read()` may return fewer bytes than a full protocol frame → you must **accumulate** bytes (often in a per-connection buffer stored as the key's attachment) until you have a complete message. It may return **0** (nothing ready) or **-1** (end of stream → peer closed: `key.cancel()` and `channel.close()`). - `write()` may consume only part of the buffer → you must **retain the remainder** and finish it later (see pitfall 2). **Fix:** treat both as loops/state machines, never one-shot. ## 4. Blocking the event loop One thread serves many channels. A blocking DB call, synchronous file read, DNS lookup, or heavy computation inside a handler **stalls every connection on that loop**. **Fix:** offload slow/blocking work to a separate worker thread pool; when it finishes, hand the result back to the loop (e.g. enqueue + `selector.wakeup()`) to write out. Never call blocking APIs from the loop thread. ## 5. ByteBuffer state mistakes A `ByteBuffer` has `position`, `limit`, `capacity`. After you `read()` *into* it, you must call **`flip()`** to switch from write-mode to read-mode before you drain/forward those bytes; before reading more you call **`clear()`** (discard all) or **`compact()`** (keep undrained bytes). Forgetting `flip()` sends garbage/zero bytes; forgetting `clear()`/`compact()` reads nothing (buffer is full). **Fix:** internalize the `read → flip → consume → compact/clear` cycle. ## 6. The epoll selector-spin bug A long-standing JDK issue (rooted in Linux epoll edge cases) can cause `select()` to **return 0 repeatedly without blocking**, pinning a CPU at 100%. The known mitigation is to detect an abnormal number of zero-return wakeups in a short window and **rebuild the selector** (open a new one, re-register all keys, swap). Netty implements exactly this; another reason teams use a framework rather than rolling their own. ## 7. Cancelled/invalid keys After `key.cancel()` or channel close, the key becomes invalid; touching it throws **`CancelledKeyException`**. The cancellation is finalized on the next `select()`. **Fix:** guard handlers with `key.isValid()` and wrap per-key handling in try/catch so one bad connection doesn't kill the loop. Always `close()` channels to free OS file descriptors (leaking fds eventually throws "Too many open files"). ## 8. Concurrency around the selector Selectors aren't built for arbitrary multi-threaded mutation. Registering or changing interest ops from another thread while the loop is in `select()` can deadlock/block. **Fix:** do mutations on the loop thread — queue the request from other threads and `selector.wakeup()` so the loop applies it. ## The meta-lesson Every item above is handled correctly, once, inside mature frameworks. For production, prefer **Netty/Vert.x**; hand-roll a selector loop mainly to learn or for the simplest cases.

  • You see one CPU core pinned at 100% and select() returning 0 in a loop. What's likely wrong?
    Either permanent OP_WRITE interest causing constant writable events, or the JDK epoll selector-spin bug. Fix the write-interest management; for the epoll bug, detect the spin and rebuild the selector (or use Netty, which does this).
  • How do you safely change a channel's interest ops from a worker thread?
    Don't mutate the selector directly from another thread; enqueue the change and call selector.wakeup() so the event-loop thread applies it on its next pass.

saying these in an interview costs you the question

  • Assuming one read() yields a complete message
  • Not handling read() == -1 (leaks the connection/fd)
  • Permanent OP_WRITE registration
  • Forgetting flip() after reading into a buffer
  • Mutating a Selector from another thread mid-select without wakeup()
  • Letting one CancelledKeyException crash the whole loop

context