A service runs on 40 application instances, each with 200 request-handling threads, in front of a Redis cache. Compare in-process request coalescing (singleflight) with a Redis-based recompute mutex for preventing duplicate work on a cache miss — and explain why you might use both.
answer
- Duplicates = instances × threads; two techniques, two factors
- Singleflight: local map key → in-flight future, free
- Redis lock: only thing that crosses processes
- Order: singleflight outside, Redis lock inside
- Share failures with current waiters, never remember them
basics
~20 sSingleflight collapses concurrent misses for the same key inside one process, so 200 threads become 1 recompute per instance — free, instant, no network. It cannot see other instances, so 40 remain. A Redis lock collapses across instances to about one. Layer them: singleflight first, Redis lock second.
solid answer
~1 min**Singleflight** keeps a per-process map from key to an in-flight computation. The first thread to miss starts the work; every other thread for that key attaches to the same future and receives the same result. It costs nothing, adds no network hop, has no TTL to tune, and cannot fail in a way that harms correctness. In the scenario given it takes 8,000 potential duplicate recomputes down to 40 — a 200× reduction from a purely local change. **A Redis mutex** (`SET lock:<key> <token> NX PX <ttl>`) is the only thing that can coordinate *across* processes, taking those 40 down to roughly 1. It costs a round trip on every miss, and it brings real design work: TTL sizing against p99 recompute, token-based release, and deciding what losers do. Use both, in that order: singleflight is free and removes the largest multiplier, so the Redis lock is contended by 40 callers rather than 8,000. Then the lock's loser strategy (serve stale, bounded poll, degrade) applies to a handful of callers per instance instead of every thread. One caveat: the coalesced result must be shared correctly — including failures, which should propagate to the current waiters but not be remembered.
code
text · 13 linesvalue = GET product:42
if value: return value
# process-local: 200 threads collapse to 1 here
return singleflight("product:42", func():
if SET lock:product:42 <token> NX PX 5000 == OK:
v = recompute()
SET product:42 v EX 300
release_lock_with_token() # Lua compare-and-delete
return v
else:
return stale_or_bounded_wait_or_degrade()
)go deeper
Know that coalescing means many threads in one process share a single computation, and that a Redis lock is what coordinates separate processes.
Do the arithmetic (instances × threads), explain why singleflight is free and process-local, and place the Redis lock inside it.
Discuss failure propagation, waiter deadlines, result immutability, and when the residual per-instance concurrency makes the distributed lock unnecessary.
Decide per workload how much duplicate recompute is acceptable, and separate deduplication (same key) from admission control (cold cache, many different keys).
## Two different multipliers Duplicate recompute count on a miss is roughly `instances × concurrent_threads_hitting_that_key`. There are two independent factors, and the two techniques attack different ones: - **Threads within a process** — handled by in-process coalescing. - **Processes across the fleet** — handled by a Redis lock. With 40 instances × 200 threads, the naive worst case is 8,000 identical recomputes. Singleflight alone: 40. Redis lock alone: ~1, but every one of the 8,000 threads pays a Redis round trip and then needs a loser strategy. Both: ~1, with only 40 lock attempts. ## How singleflight works A process-local map from cache key to a pending result: 1. Thread misses the cache, looks up the key in the in-flight map. 2. Not present → insert a placeholder (a promise/future/deferred), run the computation, then complete the placeholder and remove the entry. 3. Present → wait on the existing placeholder and return its result. All of it is a mutex or concurrent-map operation in local memory: microseconds, no network, no TTL. Most ecosystems ship it — Go's `golang.org/x/sync/singleflight`, Caffeine's `LoadingCache` (which loads a key at most once concurrently), Java's `ConcurrentHashMap.computeIfAbsent` over `CompletableFuture`, Python asyncio futures keyed by request. **Details that bite:** - **Failure sharing.** If the computation throws, all current waiters should see the failure — but the entry must be removed so the *next* request retries. Caching the failure in the in-flight map turns a transient blip into a sticky outage. If you want to suppress retries after a failure, do it deliberately with a short negative-cache entry, not by leaving a dead future in the map. - **Timeouts.** Waiters should not wait longer than their own request deadline. A hung computation must not pin every thread that asked for that key. - **Result immutability.** All waiters receive the same object; if it is mutable and one caller mutates it, the others are corrupted. Return an immutable value or a copy. - **Memory.** The map holds only in-flight entries, so it is naturally small — but the key must be the same string the cache uses, or coalescing silently does not happen. ## How the Redis mutex works, and what it costs `SET lock:<key> <token> NX PX <ttl>` gives fleet-wide coordination: exactly one process rebuilds. The costs are not in the command but in everything around it — TTL sized against p99 recompute (too short means overlap, too long means a crashed holder blocks refreshes), token-comparing release so you never delete someone else's lock, releasing on the failure path, and a policy for what losers do (return stale, poll with a deadline, degrade). Plus one extra round trip on every miss. That is worth paying when the recompute is genuinely expensive — a heavy aggregate query, a third-party API call with a quota, an ML scoring pass. It is not worth paying when the recompute is a cheap indexed lookup: 40 cheap queries are nothing, and the lock's complexity buys you little. ## Why layering wins The two compose cleanly because they operate at different scopes and neither weakens the other: ``` GET key -> hit? return singleflight(key): -> collapse this process's threads to one if SET lock:key tok NX PX -> won: recompute, SET value, release else -> lost: stale / bounded wait / degrade ``` Note the ordering: singleflight *outside*, Redis lock *inside*. That way only one thread per instance ever attempts the lock, and only one thread per instance ever waits — the other 199 are attached to the same local future and cost nothing. Reverse the order and every thread does a Redis round trip and needs its own loser handling, which is precisely the load you were avoiding. ## When singleflight alone is enough Often. Ask: is 40 concurrent identical recomputes actually a problem for the source of truth? For a database query taking 20 ms, 40 concurrent copies is a non-event; the added complexity of a distributed lock is not justified. The fleet-wide lock earns its place when the recompute is slow, rate-limited, expensive per call, or when the fleet is large enough that instance count itself is the multiplier (hundreds of pods). ## What neither of them fixes Both deduplicate work for **the same key**. A cold cache after a restart or a mass eviction produces misses for thousands of *different* keys, where every recompute is genuinely necessary. Then you need admission control instead: a bounded pool or semaphore limiting total concurrent recomputes, pre-warming of critical entries, and an explicit degradation response. And both are about the miss path — spreading TTLs so keys do not expire in lockstep, and serving stale so nobody waits, remain separate and complementary concerns.
- If the coalesced computation throws an exception, what should happen to the waiters and to the in-flight entry?Propagate the failure to the callers currently attached — they asked for a value and there isn't one — but remove the in-flight entry immediately so the next request starts a fresh attempt. Leaving a failed future in the map turns a momentary blip into a sticky outage for that key. If you want to damp retries against a failing source, add an explicit short-TTL negative cache entry rather than reusing the coalescing map for that purpose.
- When is in-process coalescing on its own sufficient, without a Redis lock?When the residual concurrency — one recompute per instance — is harmless to the source of truth. For a 20 ms indexed query across 40 pods, 40 concurrent copies is a non-event, and a distributed lock adds a round trip plus TTL tuning plus a loser strategy for no real benefit. Reach for the Redis lock when the recompute is slow, rate-limited, billed per call, or when the fleet is large enough that instance count alone overwhelms the source.
One person per office volunteers to go and ask a question (singleflight); then the offices agree that only one building sends someone (the Redis lock). Everyone else waits for the answer to come back.
saying these in an interview costs you the question
- Believing in-process coalescing prevents duplicate work across instances
- Placing the Redis lock outside the coalescing layer, so every thread makes a round trip and needs its own loser handling
- Caching the failed computation in the in-flight map, so one error keeps being returned to later requests
- Letting waiters block past their own request deadline on a hung computation
- Expecting either technique to help with a cold cache, where every miss is for a different key