skip to content

Internals

You will learn why Redis is fast and how to look inside it: the single-threaded event loop with io-threads, the RESP wire protocol, safe keyspace iteration, and the slowlog and latency tooling. Interviewers use 'why is Redis single-threaded yet fast?' as a classic depth probe.

part ofRedisoverview, primer and where to startread it →
on this pageshow

explore

questions

20

Redis is commonly described as "single-threaded". What exactly is single-threaded about a redis-server process, and what does that design buy and cost you?

level: juniorimportance: must knowfreq 78%

answer

  1. one thread owns the keyspace
  2. atomic by construction, no locks
  3. head-of-line blocking, shared p99
  4. one core per instance, scale by shards
  5. io-threads = I/O only, not execution

basics

~20 s

Command execution is single-threaded: one command at a time runs on the main event loop, so every command is atomic and no locks are needed. The cost is that one slow command stalls all clients, and one instance uses about one CPU core.

solid answer

~50 s

A redis-server process has several threads, but **only one executes commands**. The main thread runs an event loop that polls sockets with epoll/kqueue, reads a request, parses it, executes it to completion, queues the reply, then moves on. What it buys: - **Atomicity for free** - every command, Lua script and Function runs with nothing interleaved. No mutexes, no lock contention, no races on the keyspace. - **Simple internals, low latency** - an in-memory `GET` is a hash lookup, so the bottleneck is syscalls and network, not CPU. What it costs: - **Head-of-line blocking** - one O(n) command over a million elements, or a long Lua script, freezes every other client for that duration. Your p99 is a property of your worst command. - **One core per instance** - you scale with more instances/shards, not more threads. Since 6.0 `io-threads` can parallelize socket read/write, but command execution stays on the one thread.

code

text · 5 lines
text
# client A (set with 5M members)
127.0.0.1:6379> SMEMBERS huge:set     # O(N), holds the loop

# client B, at the same moment
127.0.0.1:6379> GET session:42       # waits until A finishes

go deeper

for a junior

Say it plainly: one command runs at a time, so commands are atomic and no locking is needed, and a slow command blocks everyone.

for a middle

Add the mechanics: an event loop multiplexes sockets, the thread never blocks, and command big-O directly becomes shared latency.

for a senior

Tie it to operations - p99 is shared, capacity is one core per instance, and io-threads/lazyfree move I/O and frees off the hot path without parallelizing execution.

for a principal

Frame it as a deliberate trade of vertical scaling for determinism and atomicity, and explain how sharding, pipelining and workload isolation recover the throughput you gave up.

## What is actually single-threaded A `redis-server` process is not literally one thread. It has a main thread, a few `bio` (background I/O) helper threads, and optionally `io-threads`. The single-threaded claim is about **command execution**: exactly one thread ever touches the keyspace, the dictionary mapping keys to values. Closing file descriptors, fsyncing the append-only file, and freeing very large objects are pushed to helpers; the actual mutation of data is not. The main thread runs a loop: ask the kernel which client sockets are ready, read whatever bytes arrived, parse complete commands out of the buffer, **execute each command start to finish**, append the reply to that client's output buffer, then repeat. Nothing preempts a command halfway through. ## Why the design was chosen Redis operations are almost all memory accesses. A `GET` is a hash lookup, an `LPUSH` is a pointer splice - tens to hundreds of nanoseconds. Compared to that, the cost of acquiring and releasing a lock, and the cache-line ping-pong between cores contending on shared structures, is not a rounding error: it can dominate. A single thread avoids all of it. The real bottleneck for a typical Redis workload is the network stack and the per-request syscalls, which is why pipelining (batching many commands per round trip) raises throughput far more than any threading would. ## What you get **Atomicity without effort.** Because no command interleaves with another, every command is atomic by construction, including compound ones like `INCR`, `SETNX`, `LPOP`, `GETSET` and `ZADD`. That is what makes Redis usable as a coordination primitive at all. The same guarantee extends to a whole Lua script or Function: the server runs it as one unit, so read-modify-write logic in a script is atomic without any locking. **Determinism and simplicity.** Replication ships commands (or effects) to replicas in a single well-defined order. Debugging is tractable: there is one execution order, not an interleaving. **Low tail latency when commands are small.** With only O(1) commands in the mix, per-op latency is sub-millisecond and very tight, because there is no scheduler contention on the data path. ## What it costs **Head-of-line blocking is the headline cost.** The server is a single queue with a single server. If a client issues `SMEMBERS` on a set of 5 million elements, or `LRANGE mylist 0 -1` on a huge list, or a Lua script that loops a million times, every other client waits. Latency is a shared resource: one team's O(n) command becomes another team's timeout. This is why command complexity, documented as a big-O per command, is an operational property in Redis and not just trivia, and why long-running scripts are dangerous (only `SCRIPT KILL` can stop one, and only if it has not written yet). **One core per instance.** No matter how many cores the box has, one Redis instance saturates roughly one of them for command work. Scaling out means more processes: multiple instances on one host, or Redis Cluster shards, each with its own event loop, each owning a slice of the keyspace. **Blocking is emulated, not real.** Because the thread must never block, commands like `BLPOP` and `XREAD BLOCK` do not park the thread. The client is placed on a wait list for the key and the loop continues; when another client pushes to that key, the blocked client is served. Blocking is a property of the *client*, never of the server. ## The nuance to say out loud Since Redis 6.0, `io-threads` can offload reading from and writing to sockets to extra threads, and `bio` threads plus `UNLINK`/lazyfree move large deallocations off the main thread. Neither changes the core rule: **commands still execute one at a time on the main thread**. Saying "Redis 6 made Redis multi-threaded" is the classic overstatement; the correct statement is that Redis parallelized I/O, not execution. Likewise, RDB snapshots and AOF rewrites are done by a forked *child process*, not by threads, which is another way work leaves the hot path.

  • If Redis is single-threaded, how do BLPOP and XREAD BLOCK work without stalling the server?
    They never block the thread. The command registers the client on a wait list for that key and returns control to the event loop, so other clients keep being served. When a later command pushes data to the key (or the timeout fires), the server wakes the parked client and delivers the reply. Blocking is a client-side state, not a server-side stall.
  • Does the single-threaded model mean MULTI/EXEC is unnecessary?
    No. Each individual command is atomic, but a sequence of round trips from your application is not - another client can act between them. MULTI/EXEC queues commands and runs the whole batch on the loop with nothing interleaved, and a Lua script or Function does the same for logic that needs to read a value and then decide.

A single cashier who is extremely fast. Nobody ever fights over the register, and every transaction is complete before the next starts - but the customer with 400 items freezes the whole line.

saying these in an interview costs you the question

  • Saying Redis 6 made Redis multi-threaded, so command throughput scales with cores
  • Believing a single-threaded server cannot handle many concurrent connections (it multiplexes tens of thousands)
  • Assuming slow commands only hurt the client that issued them
  • Thinking BLPOP occupies the server thread for the whole blocking period
  • Claiming atomicity requires you to add explicit locking around normal Redis commands

context

open as a page

In Redis, what actually happens when you run the KEYS command with a glob pattern against a production database with millions of keys, and what should you use instead?

level: juniorimportance: must knowfreq 70%

basics

~20 s

KEYS visits every key in the database in a single command, O(N), and Redis executes commands one at a time — so every other client waits until it finishes, potentially seconds. Use SCAN, which iterates the same keyspace in small cursor-based batches.

open as a page

Redis has a SLOWLOG facility. What exactly does it record, how do you configure and read it, and what part of a request's total time does it NOT include?

level: middleimportance: must knowfreq 60%

basics

~20 s

SLOWLOG stores the last N commands whose execution exceeded slowlog-log-slower-than microseconds. Read with SLOWLOG GET/LEN, clear with SLOWLOG RESET. It times only command execution inside the server — not network transfer, not time spent waiting in the queue behind other commands.

open as a page

Redis clients talk to the server using the REdis Serialization Protocol (RESP). Describe how a command and its reply are framed on the wire, and name the basic RESP2 data types with their prefix bytes.

level: middleimportance: must knowfreq 45%

basics

~20 s

A client sends every command as a RESP array of bulk strings; the server replies with one typed frame. RESP2 has five types identified by a first byte: + simple string, - error, : integer, $ bulk string (length-prefixed, binary safe), * array. Every frame ends with CRLF.

open as a page

The Redis SCAN command returns a cursor you pass back on the next call. What consistency guarantees does a complete SCAN iteration give you about the elements returned, and why can the same key come back twice?

level: middleimportance: must knowfreq 55%

basics

~20 s

Guarantee: any element present from the start to the end of the full iteration is returned at least once. Elements added or removed mid-iteration may or may not appear, and duplicates are possible because the hash table can rehash while you iterate. No snapshot, no server-side state.

open as a page

A Redis instance serves sub-millisecond responses most of the time but shows recurring multi-hundred-millisecond spikes. Walk through the causes you would investigate and the evidence that distinguishes them.

level: seniorimportance: must knowfreq 50%

basics

~20 s

Check four families: a blocking O(N) command or long Lua script (SLOWLOG); a fork for RDB/AOF rewrite (LATENCY fork, latest_fork_usec); memory pressure — swapping or eviction cycles at maxmemory; and bulk key expiry (LATENCY expire-cycle). Confirm the host floor with redis-cli --intrinsic-latency.

open as a page

Before blaming Redis for a slow endpoint, how would you measure an individual Redis server's own response latency and its floor on that machine, using tools shipped with Redis?

level: juniorimportance: should knowfreq 40%

basics

~20 s

Run redis-cli --latency (or --latency-history) from a client host: it sends PING in a loop and reports min/avg/max round-trip in milliseconds. Run redis-cli --intrinsic-latency <seconds> on the server itself to measure the machine's own scheduling floor with no network involved.

open as a page

Walk through one iteration of the Redis main event loop: how does a single thread serve tens of thousands of client connections, and what are file events versus time events?

level: middleimportance: should knowfreq 42%

basics

~20 s

Redis uses its own small event library (ae) over epoll on Linux or kqueue on BSD/macOS. Each iteration: run due time events, ask the kernel which sockets are readable/writable, then for each ready socket read, parse, execute and buffer the reply. Sockets are non-blocking, so the thread never waits on one client.

open as a page

Redis 6 introduced version 3 of its serialization protocol alongside the existing one. How does a client select the protocol version on a connection, and what does the newer version add?

level: middleimportance: should knowfreq 38%

basics

~20 s

A connection starts in RESP2; the client opts in by sending HELLO 3 (optionally with AUTH and SETNAME), and the server replies with a map of server info or an error on older servers. RESP3 adds typed replies — map, set, double, boolean, big number, verbatim string, null — plus out-of-band push frames and attributes.

open as a page

What do the MATCH and COUNT options of the Redis SCAN command actually control, and why can one SCAN call return zero keys even though the iteration is not finished?

level: middleimportance: should knowfreq 46%

basics

~20 s

COUNT is a hint for how much work (how many buckets) one call does — default 10 — not how many keys come back. MATCH filters the retrieved names afterwards, so it saves network, not scanning. A call can therefore return an empty batch; only cursor 0 means done.

open as a page

What does the io-threads configuration option introduced in Redis 6.0 actually parallelize, what does it deliberately not parallelize, and when is turning it on justified?

level: seniorimportance: should knowfreq 45%

basics

~20 s

io-threads offloads reading bytes from client sockets and writing replies back, plus RESP parsing, to extra threads. Command execution stays on the main thread. It helps only when the instance is bottlenecked on socket I/O with large values or very high throughput, not when commands themselves are slow.

open as a page

Redis includes a built-in latency monitoring framework exposed through the LATENCY family of commands. How does it work, how do you turn it on, and what does it show you that a log of slow commands cannot?

level: seniorimportance: should knowfreq 32%

basics

~20 s

Set latency-monitor-threshold to a millisecond value; Redis then records time-series samples for named event types (fork, expire-cycle, command, aof-fsync-always, eviction-cycle...). LATENCY LATEST shows the newest and worst per event, LATENCY HISTORY <event> the samples, LATENCY RESET clears, LATENCY DOCTOR explains them in prose.

open as a page

In the RESP protocol, most frames a client reads are replies to commands it sent. Explain what an out-of-band push frame is, why the older protocol version could not express one, and what capabilities it unlocked.

level: seniorimportance: should knowfreq 28%

basics

~20 s

A push frame (prefix >, RESP3 only) is a server-initiated message that is not a reply to any command. RESP2 had no such marker, so subscription messages looked like ordinary array replies and a subscribed connection had to be restricted to subscription commands. Push frames let a client demultiplex, so one connection can carry both commands and server notifications.

open as a page

You must remove every key matching the prefix 'session:' from a live Redis instance holding tens of millions of keys, without causing a latency incident. How do you do it, and what makes the naive approach dangerous?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Loop SCAN with MATCH session:* and a moderate COUNT, deleting each batch with a pipelined UNLINK (async free) and pausing between batches. The naive KEYS session:* piped into DEL blocks the server twice: once for the O(N) enumeration and again for one huge delete.

open as a page

Redis also offers HSCAN, SSCAN and ZSCAN for iterating inside a single collection. What do they return, and why does such a call sometimes hand back the entire collection at once with a cursor of 0?

level: seniorimportance: should knowfreq 32%

basics

~20 s

They iterate one hash, set or sorted set instead of the keyspace — HSCAN returns field/value pairs, SSCAN members, ZSCAN member/score pairs — with the same cursor rules. When the collection is small it is stored as a flat compact array, not a hash table, so Redis returns all of it in one call with cursor 0 and ignores COUNT.

open as a page

You are given a 32-core host with 256 GB of RAM to run Redis for a high-throughput workload. Given that one Redis instance executes commands on a single thread, how do you decide the deployment shape?

level: principalimportance: should knowfreq 34%

basics

~20 s

One instance uses about one core, so a big host means many instances, not one. Shard into multiple processes (Redis Cluster or per-service instances), keep each instance's dataset small enough for fast fork, replication and failover, and leave cores and RAM headroom for io-threads, bio threads and copy-on-write during snapshots.

open as a page

You can open a raw TCP connection to Redis with netcat, type `PING` followed by Enter, and get a reply — even though real clients send commands as length-prefixed arrays. What mechanism makes that work, and what are its limitations?

level: juniorimportance: nice to knowfreq 18%

basics

~20 s

Redis supports inline commands: if a line does not start with *, the server splits it on whitespace and treats the words as command arguments. It exists for interactive debugging over telnet or netcat. It cannot carry binary or spaces inside arguments, and the line has a length limit, so real clients never use it.

open as a page

Redis's wire protocol is deliberately human-readable text with CRLF line endings rather than a compact binary encoding. What are the engineering tradeoffs of that choice, and how does it still handle arbitrary binary values?

level: seniorimportance: nice to knowfreq 24%

basics

~20 s

Text framing makes the protocol trivial to implement, debug with tcpdump or netcat, and extend without versioning. Binary safety comes from length prefixes: bulk strings declare their byte count, so payloads may contain any bytes including CRLF. The cost is a few extra bytes per frame and integer-to-text conversion, which is negligible next to network round trips.

open as a page

You own a fleet of Redis instances shared by a dozen services and need a latency observability strategy rather than ad-hoc debugging. What signals would you collect, what would you alert on, and where would you accept blind spots?

level: principalimportance: nice to knowfreq 22%

basics

~20 s

Treat client-side percentiles as the SLO signal and server-side facts as explanations: scrape SLOWLOG entries by id, LATENCY LATEST per event, INFO commandstats/latencystats, fork and eviction counters. Alert on client p99 and on rates of slow entries or stall events — not on single samples. Accept that queueing victims and per-tenant attribution stay blind spots.

open as a page