A Django leaderboard is cached with cache.get_or_set(key, compute_leaderboard, 300); when the entry expires under heavy traffic the database spikes. What does get_or_set actually do, and how would you fix it?
answer
- get, then call, then add
- no lock around the callable
- every misser computes
- add as a claim, or refresh ahead
basics
~20 sDjango's get_or_set is get, call the default, add, get again, with no lock, so every request missing an expired hot key runs the expensive callable. Refresh the value ahead of expiry, or let one worker claim the rebuild with add().
solid answer
~40 sIn `BaseCache`, `get_or_set(key, default, timeout)` reads the key; on a miss it calls `default()` if callable, calls `add()` with the result, then reads the key again and returns what is stored, so concurrent callers end up returning the same winning value. Nothing stops them all computing it: fifty requests that miss in the same second run `compute_leaderboard()` fifty times, and only one `add()` wins. Django has no built-in lock or early-refresh option for this. Practical fixes: write the leaderboard from the code that changes it, or from a scheduled job, with `timeout=None`, so readers never miss; or have one request claim the rebuild with `cache.add("leaderboard:rebuild", 1, 30)`, which is atomic on Redis and Memcached, while others serve a stale copy kept under a second key.
code
python · 23 linesfrom django.core.cache import cache
FRESH = "leaderboard:weekly"
STALE = "leaderboard:weekly:stale"
CLAIM = "leaderboard:weekly:rebuild"
def get_leaderboard():
rows = cache.get(FRESH)
if rows is not None:
return rows
if cache.add(CLAIM, 1, timeout=30):
# This request won the claim: rebuild once.
try:
rows = compute_leaderboard()
cache.set(FRESH, rows, timeout=300)
cache.set(STALE, rows, timeout=3600)
finally:
cache.delete(CLAIM)
return rows
# Someone else is rebuilding: serve the stale copy if there is one.
stale = cache.get(STALE)
return stale if stale is not None else compute_leaderboard()go deeper
Recall that get_or_set reads a key and, on a miss, stores and returns a default that may be a callable.
Explain the get, call, add, get sequence and why callers agree on the value but may all compute it.
Diagnose the expiry spike and fix it with refresh-ahead writes, an add() claim on an atomic backend, or a stale fallback.
Decide which hot values deserve precomputation and who owns refreshing them, rather than relying on lazy fills.
## What get_or_set does, line by line `cache.get_or_set(key, default, timeout=DEFAULT_TIMEOUT, version=None)` is implemented once in `django.core.cache.backends.base.BaseCache` and inherited by every built-in backend: 1. `get(key)` with an internal sentinel. A hit returns immediately. 2. On a miss, if `default` is callable, **call it**. Otherwise use it as the value. 3. `add(key, value, timeout)`: store it only if nobody else has in the meantime. 4. `get(key)` again, and return **what is stored**, which may be another caller's value. Step 4 exists, as the source comment says, to avoid a race when another caller adds a value between the first `get()` and the `add()`. It makes callers **agree on the value**; it does nothing to stop them **computing** it. ## Why the database spikes With a popular leaderboard and a 300-second timeout: - At second 300 the entry expires. - Every request arriving in the next moment runs step 1 and misses. - Each one runs `compute_leaderboard()`, the expensive aggregate query, in parallel. - One `add()` wins; the rest are discarded, but the queries have already run. There is **no lock, no single-flight and no early refresh** inside Django's cache API. The general theory of this failure mode is a system-design topic; the Django-specific point is that `get_or_set()` offers no protection, and passing a callable only delays the work until a miss, it does not serialise it. ## Fixes using Django's own API | Approach | How in Django | Trade-off | |---|---|---| | Refresh ahead | A scheduled job or the write path calls `cache.set(key, rows, None)` | Readers never miss; needs a trigger that always runs | | Claim the rebuild | `cache.add("leaderboard:rebuild", 1, 30)`; only the `True` caller recomputes | Others must have something to serve meanwhile | | Stale copy | Keep `leaderboard:stale` with a longer timeout; serve it while one caller rebuilds | Readers may see slightly old data | | Jittered timeouts | Vary `timeout` per key so related entries do not expire together | Reduces, does not remove, the spike | The claim pattern relies on `add()` being atomic on the backend: - **`RedisCache`**: `add()` is a single `SET ... NX` command. - **Memcached backends**: `add` is a native server operation. - **`LocMemCache`**: atomic within one process only, under its lock. - **`FileBasedCache`**: `add()` checks and then writes, so two processes can both succeed. ## Invalidating the leaderboard When a match result changes the rankings: - **Delete** the key (`cache.delete("leaderboard:weekly")`) and let the next reader rebuild, which reintroduces the miss; or - **Overwrite** it immediately from the write path with `cache.set(...)`, which keeps readers on a hit. Overwriting from the write path is usually better for a hot key, provided the write path can afford the computation. ## Verifying the fix Before and after changing the code, measure rather than guess: 1. **Reproduce the spike**: in a staging environment, fire a burst of concurrent requests just after the key expires and count how many times `compute_leaderboard()` runs, for example with a log line or a counter. 2. **Check the backend**: the claim pattern only works where `add()` is atomic across processes, so confirm the alias is `RedisCache` or a Memcached backend and not `LocMemCache` or `FileBasedCache`. 3. **Watch the fallbacks**: log how often requests serve the stale copy and how often they fall through to computing directly; the second number should be close to zero once the stale key is warm. 4. **Test the failure path**: if `compute_leaderboard()` raises, the `finally` block must release the claim, otherwise nobody rebuilds until the claim's timeout passes. ## Other get_or_set details - The **returned value** is what is stored, so a caller can receive a leaderboard computed by another worker a moment earlier. - With **`timeout=0`**, `add()` stores nothing and `get_or_set()` returns the computed default every time, silently disabling the cache. - A callable returning **`None`** is stored as `None`, and later calls treat that as a hit, because `get_or_set()` compares against an internal sentinel rather than `None`. - `aget_or_set()` follows the same steps with awaited calls.
- Two workers miss at the same time and both call get_or_set with a callable. Which value does each return?Both run the callable, but only the first `add()` stores its result. Each then calls `get()` again and returns the stored value, so both return the winner's leaderboard. The loser's computation is discarded, which is why the database still sees two queries.
- Why is the add()-based claim unsafe on FileBasedCache?`FileBasedCache.add()` checks whether the key exists and then writes the file as two separate steps, so two processes can both see no file and both write, each believing it won. On Redis the claim is one `SET NX` and on Memcached one native `add`, so exactly one caller succeeds.
- What does get_or_set return with timeout=0?The freshly computed default, every time. `add()` with a timeout of 0 stores nothing (or stores and immediately removes it), so the second `get()` misses and returns the default passed to it. The cache is silently bypassed on every call.
saying these in an interview costs you the question
- get_or_set holds a lock so only one caller runs the callable.
- Passing a callable to get_or_set prevents concurrent recomputation.
- get_or_set always returns the value this caller computed.
- cache.add() is atomic across processes on every Django backend.
- A shorter timeout reduces the database spike when the key expires.