What does the io-threads configuration option introduced in Redis 6.0 actually parallelize, what does it deliberately not parallelize, and when is turning it on justified?
answer
- io-threads = socket read/write + RESP parse
- execution stays on main thread
- io-threads-do-reads no by default
- helps I/O-bound, big values, spare cores
- not a fix for O(n) commands or Lua
basics
~20 sio-threads offloads reading bytes from client sockets and writing replies back, plus RESP parsing, to extra threads. Command execution stays on the main thread. It helps only when the instance is bottlenecked on socket I/O with large values or very high throughput, not when commands themselves are slow.
solid answer
~50 s`io-threads N` (Redis 6.0+) creates N-1 helper threads that the main thread hands socket work to in batches. They perform the `read()`/`write()` syscalls and RESP protocol parsing and reply encoding. The main thread then **executes all commands itself, one at a time**, exactly as before, so atomicity and ordering are unchanged. By default only writes are offloaded; `io-threads-do-reads yes` extends it to reads and parsing. It is worth enabling when profiling shows the main thread burning its time in `read`/`write` and protocol work - typically very high ops/sec, large values, or many clients on a fat NIC - and you still have spare cores. It does **nothing** for a workload limited by expensive O(n) commands or Lua scripts, and on a small instance the coordination overhead can make things slightly worse. Redis's own guidance is not to enable it unless you have measured the bottleneck, and to leave at least one core free.
code
text · 7 lines# redis.conf - leave cores free for OS, bio threads and forks
io-threads 4
io-threads-do-reads yes
# check whether the bottleneck is really I/O before enabling
redis-cli INFO commandstats | head
redis-cli SLOWLOG GET 10 # empty + pegged CPU => likely I/O boundgo deeper
Know the one-liner: io-threads parallelizes socket I/O and protocol work only; commands still run one at a time.
Add the config detail - writes offloaded by default, reads only with io-threads-do-reads - and that atomicity is unaffected.
Lead with how you decide: evidence that the main thread is I/O-bound, spare cores, large payloads; and name the cases where it is useless.
Position it against the real scaling lever - sharding - and reason about core budgeting across the main thread, io-threads, bio threads and fork children.
## The problem io-threads was built for On a busy Redis instance serving simple commands, the main thread often spends more time moving bytes than doing keyspace work. A `GET` that returns a 20 KB value costs a hash lookup (nanoseconds) plus a `write()` of 20 KB plus protocol encoding (microseconds). Profiling such instances showed the event loop dominated by socket syscalls, `memcpy`, and RESP parsing/serialization - work that is *embarrassingly parallel* because each client's bytes are independent. ## What Redis 6.0 did `io-threads N` starts N-1 additional threads (the main thread counts as one of the N). The flow becomes: 1. The main thread collects the set of clients with pending reads or pending writes. 2. It **partitions that list round-robin across the io-threads**, then signals them. 3. Each io-thread performs its clients' `read()` (and RESP parsing into command structures) or reply encoding and `write()`. 4. The main thread waits for all io-threads to finish the batch (a spin/barrier), then **executes the parsed commands itself, serially**. That last step is the crux: the io-threads never touch the keyspace. There is no locking added anywhere, because no shared data structure is mutated concurrently. Command order is still whatever the main thread decides, replication is still a single ordered stream, and every command and script remains atomic. ## Configuration and defaults ``` io-threads 1 # default: disabled (main thread only) io-threads-do-reads no # default: offload writes only ``` By default, even with `io-threads` set, only *writes* (reply encoding + `write()`) are offloaded. Reads and parsing move off only with `io-threads-do-reads yes`, which historically gave a smaller and more workload-dependent benefit. `io-threads` cannot be changed at runtime in a meaningful way on older versions - treat it as a restart-level config and verify on your build. ## When it actually helps Enable it only after measuring. The signature of an I/O-bound instance is: high ops/sec or high bytes/sec, main thread at ~100% CPU, but `SLOWLOG` empty and command complexity low - i.e. the CPU is not going into command execution. Large values (kilobytes, not bytes) amplify the effect because encoding and copying scale with payload size. Concretely it is justified when: - Throughput is capped and the box has idle cores. - Values are large, or the client count and connection churn are high. - Commands are O(1) - the work is genuinely in the pipes. Redis's guidance is to set `io-threads` to fewer than the number of cores (leaving headroom for the OS, `bio` threads, and any forked child), and typically not above ~8, since coordination overhead grows. ## When it does nothing, or hurts - **CPU-bound commands.** If your latency comes from `SORT`, big `LRANGE`s, `ZRANGEBYSCORE` over huge ranges, or Lua, io-threads cannot help - that work is on the main thread by design. - **Low throughput.** The batching barrier costs a spin per iteration; on a lightly loaded instance you pay overhead for nothing and can see slightly worse latency. - **Small values and pipelining already in use.** Pipelining amortizes syscalls, which is the same win from a different direction, and is usually the cheaper fix. ## What to say about "Redis is multi-threaded now" The honest framing: Redis parallelized **I/O**, not **execution**. Redis 6.0 introduced io-threads; later Redis and Valkey releases have refined the same idea - deeper async I/O, better thread affinity - but on a given shard, commands still execute one at a time. Anyone claiming that raising `io-threads` multiplies command throughput on CPU-bound workloads is wrong. The alternative for CPU-bound scaling remains what it always was: shard the keyspace across more instances (Redis Cluster or multiple instances per host), each with its own event loop and its own core. ## Operational cautions io-threads spin while waiting for work in some versions, so CPU accounting can look alarmingly high even when idle - pin and size accordingly, and never allocate more io-threads than free cores, or you will contend with the main thread and make latency worse than the baseline you were trying to fix.
- Does enabling io-threads change atomicity, command ordering, or replication semantics?No. The io-threads only read bytes, parse RESP and write encoded replies; they never touch the keyspace. The main thread still executes every parsed command serially, so single-command atomicity, MULTI/EXEC, and script atomicity are unchanged, and the replication stream keeps a single deterministic order.
- Your latency is bad and SLOWLOG is full of O(n) commands. Would raising io-threads help?No - that is the case io-threads cannot address, because the time is spent executing commands on the main thread. The fixes are to change the access pattern (bounded ranges, SCAN-style iteration, cursor-based reads), to move heavy work out of Redis, or to shard so each instance handles less of that work.
- How many io-threads should you configure on a 16-core host running one Redis instance?Fewer than the cores you actually have free - commonly 4 to 8 - because io-threads can spin and will contend with the main thread, bio threads, and any forked child during a BGSAVE. Start low, measure ops/sec and p99 at the same offered load, and stop increasing when the gain flattens.
saying these in an interview costs you the question
- Claiming io-threads makes command execution parallel or multiplies CPU-bound throughput
- Setting io-threads equal to or above the core count
- Enabling it as a first response to latency without checking whether the bottleneck is I/O
- Assuming reads are offloaded by default (they are not, without io-threads-do-reads)
- Thinking io-threads removes the need to shard for CPU scaling