skip to content

You need to remove a Redis key holding a 20-million-element set. Compare DEL and UNLINK, and explain what Redis's bio background threads and the lazyfree settings do here.

level: middleimportance: must knowfreq 52%

answer

  1. DEL = unlink + free, both on main thread
  2. UNLINK = O(1) unlink, free on bio lazyfree thread
  3. bio threads: close file, AOF fsync, lazyfree
  4. lazyfree-lazy-{eviction,expire,server-del,user-del}
  5. FLUSHALL ASYNC, replica-lazy-flush

basics

~20 s

DEL frees every element on the main thread, so removing a huge collection blocks all clients. UNLINK removes the key from the keyspace immediately and hands the expensive free to a background lazyfree thread. The lazyfree-lazy-* settings apply the same behaviour automatically to eviction, expiry and overwrites.

solid answer

~50 s

Both commands make the key disappear atomically; they differ in **who pays for the memory reclamation**. - `DEL` unlinks the key and frees all 20 million elements synchronously on the main thread. That is O(N) with a large constant - easily hundreds of milliseconds - and every other client is stalled for it. - `UNLINK` removes the key from the dictionary in O(1) so it is instantly invisible to all clients, then estimates the free cost. If it is significant, the object is pushed to the **bio lazyfree thread**, which does the deallocation off the event loop. Small objects are still freed inline, because a thread handoff would cost more than the free. The same mechanism is available implicitly through config: `lazyfree-lazy-eviction`, `lazyfree-lazy-expire`, `lazyfree-lazy-server-del`, `replica-lazy-flush`, and `lazyfree-lazy-user-del` (which makes plain `DEL` behave like `UNLINK`). `FLUSHALL ASYNC` / `FLUSHDB ASYNC` are the same idea for a whole database. On any instance with large aggregate values, prefer `UNLINK` and turn the lazyfree flags on.

code

text · 9 lines
text
127.0.0.1:6379> SCARD huge:set
(integer) 20000000
127.0.0.1:6379> UNLINK huge:set     # O(1) on the event loop
(integer) 1

# make implicit deletes lazy too
127.0.0.1:6379> CONFIG SET lazyfree-lazy-expire yes
127.0.0.1:6379> CONFIG SET lazyfree-lazy-eviction yes
127.0.0.1:6379> CONFIG SET lazyfree-lazy-server-del yes

go deeper

for a junior

Know that DEL frees everything inline while UNLINK removes the key instantly and frees in the background, and that UNLINK is the safe choice for big collections.

for a middle

Add the effort-estimate detail, name the bio lazyfree thread, and list the lazyfree-lazy-* settings including their default of no.

for a senior

Frame it as latency management: implicit delete paths (expiry, eviction, overwrite, replica flush) are the ones that surprise you, so configure them, and treat huge keys as the root cause.

for a principal

Argue about data modelling - a design that produces twenty-million-element keys creates hazards for replication, resharding and failover, not just deletion - and set lazyfree defaults as a fleet policy.

## Why deleting can be slower than writing A Redis key holding a set, hash, list, sorted set or stream is not one allocation. A 20-million-element set is 20 million entries plus dictionary/listpack structures, each of which the allocator must return. `free()` per element is cheap individually but the total is easily hundreds of milliseconds, and it all runs on the single command thread. So the counter-intuitive fact is that **removing one key can be the slowest operation your Redis instance ever performs**, and it stalls every client while it happens. ## DEL `DEL key` removes the key from the keyspace *and* reclaims all its memory, synchronously, before returning. Complexity is O(1) for a string, O(N) in the number of elements for aggregate types. On a huge collection this is a classic latency spike: a monitoring graph shows a single 400 ms stall and nothing in the command mix looks expensive, because everyone reads `DEL` as "cheap". ## UNLINK `UNLINK key` (Redis 4.0+) splits the two jobs: 1. **Unlink** - remove the key from the database dictionary. O(1). From this instant, every client sees the key as gone; `EXISTS` returns 0. The command is atomic and its visible semantics are identical to `DEL`. 2. **Free** - Redis computes an *effort estimate* for reclaiming the object (roughly, how many allocations it involves). If the estimate exceeds a small threshold and the object is not shared, the object pointer is queued to the **lazyfree** background job; otherwise it is freed inline, because for a short string the queue handoff would be pure overhead. The reply is the number of keys unlinked, same as `DEL`. Memory does not return to the allocator instantly - `INFO memory` may lag by a moment - which is the only observable difference besides latency. ## bio threads Redis has a small pool of **bio** (background I/O) threads created at startup, each servicing a job queue for work that must never touch the event loop: - **Close file** - `close()` on a large file (for example an old AOF) can block on some filesystems, so it is queued. - **AOF fsync** - with `appendfsync everysec`, the `fsync()` is performed by a bio thread so a slow disk does not stall commands. - **Lazy free** - deallocating objects handed over by `UNLINK` and the lazyfree settings. They are true threads inside the process, not forked children, and they never touch the keyspace - by the time an object reaches the lazyfree thread it is already unreachable from the dictionary, so there is nothing to synchronize. ## The lazyfree settings Explicit `UNLINK` only helps where your code calls it. Several internal paths also delete objects, and each has a flag (in `redis.conf`): ``` lazyfree-lazy-eviction no # freeing when maxmemory eviction removes a key lazyfree-lazy-expire no # freeing when a TTL expires a key lazyfree-lazy-server-del no # freeing the old value on implicit deletes (e.g. SET over a big key, RENAME onto an existing key) replica-lazy-flush no # freeing the old dataset when a replica does a full resync lazyfree-lazy-user-del no # make an explicit DEL behave like UNLINK ``` These ship as `no` in the stock config for historical compatibility; on any instance holding large aggregate values, turning them on is close to a default best practice. `lazyfree-lazy-expire` in particular matters when big keys carry TTLs: otherwise an expiry fires during `serverCron` and you get an unexplained stall with no client command to blame. `replica-lazy-flush` matters because a replica dropping a multi-gigabyte dataset before loading a fresh RDB is a long synchronous free at the worst possible time. `FLUSHALL ASYNC` and `FLUSHDB ASYNC` apply the same trick at database scale: swap in an empty dictionary and let bio threads reclaim the old one. ## How to answer the original question Use `UNLINK` (or set `lazyfree-lazy-user-del yes` and keep calling `DEL`). Better still, avoid needing it: keys that grow to twenty million elements are big-key hazards for every operation, not just deletion - they also make cluster resharding, replication and `DEBUG` inspection painful. If the value must be that large, shard it across many keys so that no single free, dump or migrate is ever huge.

  • Is UNLINK always faster than DEL?
    No. For a small value - a short string or a ten-field hash - Redis frees it inline anyway, because queueing the object to a background thread costs more than the free itself. UNLINK wins specifically on large aggregate values, where the deallocation is thousands or millions of frees.
  • A key with a TTL holds a two-gigabyte hash. What happens at expiry, and how do you avoid a stall?
    With the default settings, the key is expired by active expiry in serverCron or lazily on access, and its memory is freed synchronously on the main thread - a multi-hundred-millisecond stall attributable to no client command. Setting lazyfree-lazy-expire yes moves that free to a bio thread. The deeper fix is not to keep single keys that large.
  • Do bio threads ever read or write the keyspace?
    No. They handle only work on data already detached from the keyspace or unrelated to it: closing file descriptors, fsyncing the AOF, and freeing objects that have already been unlinked from the dictionary. That is why no locking is required around them.

DEL is carrying every box out of the warehouse yourself before you can serve the next customer. UNLINK is taking the sign off the door and letting the night crew haul the boxes.

saying these in an interview costs you the question

  • Believing DEL is O(1) for every type because 'it just removes a key'
  • Thinking UNLINK makes the key visible for a while longer (it is removed atomically and immediately)
  • Assuming the lazyfree-lazy-* options are on by default
  • Attributing an expiry-driven latency spike to a client command because no slow command appears in SLOWLOG
  • Saying bio threads execute commands or need locks against the main thread

context